返回列表
🧠 阿头学 · 🪞 Uota学

为持久化智能体设计 Grok Bot 界面范式

Grok Bot 通过“以实体替代会话”的界面重构与“能力共享/上下文隔离”的架构设计,试图将 AI 从问答工具升级为可委派的数字同事,但其拟人化包装严重掩盖了底层模型可靠性不足与算力成本失控的核心风险。
打开原文 ↗

2026-09-04 原文链接 ↗
阅读简报
双语对照
完整翻译
原文
讨论归档

核心观点

  • 交互锚点从会话转向实体:持久化责任要求产品必须围绕 Bot 身份组织,因为临时聊天记录无法承载跨周期的上下文积累与任务问责。
  • 状态可见性采用三级阶梯:通过头像微表情与“状态-预览-接管”分层暴露,产品有效平衡了黑盒焦虑与过度干预,将主动监控降维为潜意识感知。
  • 多智能体架构强制隔离上下文:账户级共享工具技能而 Bot 级隔离记忆,直接切断了全局记忆污染路径,这是角色化协作可扩展的唯一工程解法。
  • 信息输出必须异构化:将内联卡片与结构化 UI 直接嵌入对话时间线,彻底消除了用户二次整理数据的认知摩擦,使交付形态本身成为答案。

跟我们的关联

  • 对 ATou 意味着产品评审标准必须从“功能堆砌”转向“委派体验”,下一步需强制引入“是否增加管理事务”过滤器,直接砍掉暴露底层元数据的冗余面板。
  • 对 Neta 意味着多智能体调度架构必须采纳“能力共享/上下文隔离”原则,下一步需重构记忆池,按角色边界严格隔离长期上下文以阻断幻觉级联。
  • 对 Uota 意味着交互组件库需补充渐进式透明度控件,下一步应制定基于头像生命周期的微交互规范,用潜意识状态反馈替代传统进度条。

讨论引子

  • 当底层模型仍无法保证长程任务零幻觉时,“拟人化同事”界面是否会因过度承诺而反噬用户信任?
  • “能力共享、上下文隔离”架构在跨 Bot 复杂协作中,如何界定信息同步的触发阈值以避免角色盲区?
  • 产品边界设定(单账户 50 Bot/单群聊 6 Bot)是认知负荷的最优解,还是当前算力成本妥协的临时遮羞布?

为持久化智能体的世界设计 Grok Bot

我们如何为超越单一会话持久存在的智能体设计 Grok Bot——从聊天记录到 Bot 名册、在线状态、Bot 自己的计算机,以及无需提示词即可开始的工作。

当我们开始设计 Grok Bot 时,核心问题之一是界面应如何塑造用户与智能体之间的关系。大多数 AI 界面都围绕用户操作的聊天会话组织。每个会话以设置开始,在用户的注视下展开,并在对话停止时结束。

我们想要设计一种超越任何单一会话持久存在并能独立承担责任的智能体。这意味着需要重新考虑界面的一些基本对象和信号,包括侧边栏应该包含什么、智能体如何展示进度,以及它的工作何时应该变得可见。

搜索

Kenny

7:34 PM

需要你对周五全员大会演示文稿的确认。

Justin

起草了 8 份介绍信——放在 CRM 里等你发送。

Luke

7:34 PM

收件箱有 3 封邮件。今天有两封需要回复。

网站发布

11:18 AM

John:预发环境结账流程正常,已关闭 3 个 bug。

John

昨天

复现了结账崩溃问题。详情已写在工单里。

Keith

Acme 情况不稳定。起草了一份周四汇报。

Tyler

9:04 AM

已录入 14 张收据。仍缺你周二的 Uber 记录。

Manuel

2:20 PM

发布文章已上线。前 200 次展示。

Jenny

周二

SoMa 区有 3 处房源。Folsom 的两居室就是那个。

Chang

10:12 AM

筛选了 3 份。跳过了一份已在你的 ATS 中的。

插件

Peng Zheng

Kenny

9:41 AM

嘿,今天设计同步会的主要结论是什么?

团队花了大部分时间审查新的引导流程。

最大的讨论围绕空状态展开。Sarah 觉得当前的插画不符合新的品牌方向,屋里大多数人都同意。关于进度指示器应该放在页头还是侧边栏,大家也进行了长时间的反复讨论。

不过整体情绪很积极。大多数人认为该流程已接近准备好进行更广泛的审查。

有人提起第三季度路线图吗?

是的,提了两次。Marcus 说路线图审查现在预计在八月第一周进行,Priya 问引导工作会在那之前还是之后完成。会上没有确定确切日期。

给 Kenny 发消息

Kenny 的屏幕

例行程序

早间简报

每天上午 8:00

收件箱清理

工作日下午 6:00

每周团队更新

已暂停

重新思考基本原语

AI 产品在短时间内积累了大量词汇。聊天、会话、模型、上下文窗口、记忆、系统提示词、项目、技能、连接器、智能体、工具、沙盒、权限和自动化都描述了这些系统的真实组成部分。

但是,将每一个都作为单独的产品概念暴露出来,会让用户去理解超出他们需求的内容。我们首先提出的问题是:一个人实际上需要哪些概念才能与智能体协作。

我们不断回到这五个概念:

  • Bot 是持久化的智能体,拥有自己的身份、记忆、运行时和工具。

  • 聊天是与 Bot 协作的对话界面。

  • 提示词为 Bot 提供上下文或指令。它们可以被使用一次、保存为技能,或作为例行程序自动触发。

  • 工具让 Bot 能够通过软件、API、连接器、命令行或计算机使用来访问信息并采取行动。

  • 产物是 Bot 创建或修改的文档、设计、代码、数据和其他持久化输出。

其他一切都可以隐藏在界面之下,直到用户有理由关心它们。下一个问题是,这五个对象中哪一个应该用来组织产品。

从聊天记录到 Bot 名册

聊天是一次性的。我们开始一段对话是为了解决一个问题。它被推到侧边栏下方。一周后,我们又开始了一段新的对话。你很少会翻看最近五条以外的记录。

当交互的单位是一个问题时,这种行为是完全合理的。但当交互的另一端应该了解你、记住之前的工作并随着时间的推移承担责任时,这就变得奇怪了。

因此,Grok Bot 中的主要对象是 Bot,而不是对话。Bot 有名字。它有头像和标题。它记得与你的对话。它有自己的计算机和工具。当你明天回来时,你回到的是同一个 Bot。

Project Acme

在周五通话后给 Acme 起草一封跟进邮件

重写定价单页

我在安全审查中应该问什么?

根据我的通话记录构建支持者地图

用尖锐的反对意见练习演示

帮我比较这三个竞争对手的演示文稿

显示更多

Kenny

7:34 PM

需要你对周五全员大会演示文稿的确认。

Justin

起草了 8 份介绍信——放在 CRM 里等你发送。

John

昨天

复现了结账崩溃问题。详情已写在工单里。

Keith

Acme 情况不稳定。起草了一份周四汇报。

在线状态作为界面

一旦 Bot 成为你随着时间维护的东西,而不是你开启的会话,Bot 在产品中出现的方式就必须同时回答三个问题:

  • 这是谁?

  • 他们在做什么?

  • 我需要了解多少?

这是谁

名册只有在能被快速扫视时才起作用。随着名册增长,我们不希望人们每次打开产品时都要阅读每个名字。他们应该几乎在余光中就能从头像认出一个 Bot。

Bot 头像视觉风格探索,作者:Kenny Kuh 和 Peng Zheng

同时,我们希望保持头像足够一致,使其看起来像一个系统。我们研究了插画、动画、游戏和界面设计中的角色系统,探索了从首字母和表情符号到像素艺术、水彩、黏土风、Noritake 风格线条画、剪影和身份图标等各种风格。

大多数方法只能更好地解决问题的一面。水彩和黏土风格赋予了单个 Bot 丰富的个性,但在侧边栏的尺寸下细节过多。更简单的系统在界面中显得更自然,但往往让 Bot 看起来可以互换。

我们最终确定的系统保持了基本结构的一致性,使用简单的形状和富有表现力的眼睛,然后通过受控的变化和配饰引入差异性。每个 Bot 都能一眼认出,而不会显得来自不同的视觉世界。

它们在做什么

一旦头像成为 Bot 的身份标识,它自然也就成了展示状态的理想位置。一个 Bot 可能处于空闲、思考、工作、等待、阻塞或完成状态。我们本可以用单独的指示器来表示每种状态,但这会给用户增加一层需要解读的 UI。

相反,我们探索了头像本身能够承载多少生命周期。

休息时,Bot 平静且略带好奇。当任务到来时,它会确认任务。工作开始时,它便进入状态。当它在等待或需要帮助时,动作会再次改变,一旦工作完成就会平静下来。现在,头像不仅展示了这是哪个 Bot,还展示了 Bot 正在做什么。

空闲工作等待阻塞思考完成

头像动作系统,作者:Benji Taylor

我需要了解多少

一个相关的设计问题是,应该展示多少 Bot 的执行过程。一种方法是使用标准的“三个动画点”,但这提供的信息太少了,使得用户很难判断 Bot 是在工作还是卡住了。

正在搜索网络

  • 环境就绪 387ms

  • 编辑了 math.ts +14 −10

  • 运行了聚焦测试 npm test

  • 运行了类型检查 npm run

  • 简短思考

  • 搜索了代码 “toFixed”

  • 读取了 AGENTS.md

  • 编辑了 math.test.ts +6 −2

  • 运行了全套测试 212 个通过

  • 提交了修复:clamp NaN

  • 环境就绪 387ms

  • 编辑了 math.ts +14 −10

  • 运行了聚焦测试 npm test

  • 运行了类型检查 npm run

  • 简短思考

  • 搜索了代码 “toFixed”

  • 读取了 AGENTS.md

  • 编辑了 math.test.ts +6 −2

  • 运行了全套测试 212 个通过

  • 提交了修复:clamp NaN

我们也尝试过展示 Bot 当前动作的简短文字描述,但一旦人们能看到一个步骤,他们就想看到其余的步骤。用户研究表明,他们要求提供这些细节,主要是为了确信 Bot 仍在工作且方向正确。

在最终设计中,头像的动作通过展示 Bot 处于活跃状态,提供了第一重安心感。如果有人想检查它在做什么,可以将鼠标悬停在上面查看其当前动作。

它们的电脑,不是你的

每个 Bot 都有自己的电脑,它可以用它来浏览网页、处理文件和运行软件。这带来了另一个界面问题。这台电脑应该有多显眼,用户什么时候应该能够控制它?

我们探索了四种布局:

  • 浮动窗口:保持电脑易于访问,但遮挡了对话。

  • 并排显示:使工作过程持续可见,并鼓励用户观看。

  • 模态窗口:使查看变得容易,但将 Bot 的工作区视为临时中断。

  • 全屏:给电脑提供了充足的空间,但完全取代了对话。

我们让电脑越显眼,产品就越鼓励用户去监督它。我们决定它应该保持为 Bot 的工作区,界面根据用户的需求提供不同级别的访问权限。

最终设计有三个级别,允许用户进入 Bot 的工作区,而不必被卷入操作中:

  • 状态:当电脑处于活动状态时,标题栏图标变为紫色。

  • 预览:打开它会显示一个固定的侧边面板,用户可以在不离开对话的情况下跟进工作。

  • 接管:当 Bot 需要帮助时,用户可以全屏打开电脑,接管控制权,然后再交还给它。

我们还设计了随时间变化的壁纸,在早晨变亮,在夜晚变暗。这个细节赋予了 Bot 的电脑自己的时间感,使其感觉上独立于用户的桌面。

动态壁纸由 Kenny Kuh 和 Luke Barker 创作。

这更像是与一位同事共事,而不是操作一台远程机器。你能看出他们正在工作,在需要上下文(Context)时可以瞥一眼他们的屏幕,并在有事情需要你帮忙时坐下来处理。

信息的形态

Grok Bot 的早期版本几乎对每个请求都用散文(Prose)来回复。它会描述未来五天的天气预报,而不是直接展示出来;它会口述一组任务,而不是将它们作为看板(Board)排列出来。然后用户不得不重新组织这些答案。这促使我们将回复的形式视为答案的一部分。

为了支持这一点,我们在 Grok Bot 中内置了内联卡片(Inline cards)和小部件(Widgets)。当散文适合传达信息时,Bot 可以用散文回复;当不适合时,则使用结构化 UI(Structured UI)。

新邮件

准备发送

发件人 peng@grokbot.app

收件人 sarah@acme.com

主题 将周五的设计评审改到下午 2 点

你好 Sarah,

我们能把周五的设计评审从上午 11 点改到下午 2 点吗?刚接到一个客户电话,我不想仓促进行我们的讨论。

谢谢, Peng

发送邮件 丢弃

内联聊天小部件,由 Peng Zheng 提供

同样的原则也适用于操作(Actions)。当 Bot 创建一个例行程序(Routine)、更改设置或向另一个 Bot 发送消息时,该事件可以直接显示在对话记录(Transcript)中。如果有更多需要查看的内容,用户可以将其展开。

上午 9:41

早上好!你能帮我跟所有人确认一下进度吗?

正在处理——现在就向团队询问状态

与 Kenny、Tyler 和 Jenny 的 6 条消息

一切顺利:Kenny 发布了落地页,Tyler 发送了本月的发票,Jenny 预订了下周的面试。没有阻碍。

太棒了,你每天早上都能这样做吗?

已创建例行程序 晨间简报

完成,你的晨间简报每天早上 9:00 会发到这里

结果是一个异构的对话记录(Heterogeneous transcript),其中对话、系统事件、交互对象和可视化内容共享同一个时间线。

组织智能

一旦人们创建了多个 Bot,产品就必须组织好这些 Bot 如何协同工作。我们需要决定哪些上下文(Context)应该属于每个角色,当 Bot 的工作重叠时它们应该如何共享上下文,以及如何协调它们而不把用户变成调度员(Dispatcher)。

随着人们创建的 Bot 越来越多,我们看到了一种答案。有些人创建了一个“幕僚长”Bot(Chief of Staff Bot)来负责协调几位专家。他们可以向一个 Bot 下达指令,而不必逐个检查并亲自分配每项任务。

赋予 Bot 不同的角色也迫使我们决定每个角色应该知道什么。一个法务 Bot 可能需要了解正在进行的纠纷的历史,而一个财务 Bot 可能需要数年的财务记录。将这些历史合并到一个庞大的记忆(Memory)中,会使向每个 Bot 提供与其工作相关的信息变得更加困难。

因此,在 Grok Bot 中,能力(Capabilities)和上下文遵循不同的边界。工具(Tools)和技能(Skills)存在于账户级别,因为许多 Bot 可能都需要浏览网页、处理文档或发送电子邮件。记忆和例行程序(Routines)则属于特定的 Bot,因为它们反映了该特定角色随着时间推移所掌握的知识和所做的事情。换言之,能力可以被广泛共享,而上下文则保留在需要它的角色身边。

有些工作会跨越这些角色边界。群聊为项目或团队提供了共享的上下文,同时允许每个 Bot 保留其专门的记忆。设计师、工程师、项目经理(PM)和数据科学家可以在同一个对话中工作,相互交接工作,并共享项目所需的内容。

我们曾考虑添加仪表盘(Dashboards)、任务分配看板(Assignment boards)和明确的交接控制(Handoff controls)来管理这些群组。但每一种都给用户增加了更多的协调工作。相反

到了项目后期,许多设计工作都涉及做减法。我们移除了窗口和面板控件、计算机视图选项,以及智能体(Agent)元数据。我们还设定了实际限制:每个账户大约 50 个 Bot,每个群聊 6 个 Bot。每一个决策最终都归结为同一个问题:这究竟是帮助人们更好地委派任务,还是又给他们增加了一项需要管理的事务?

随着模型的不断进步,操作 AI 与向同事委派任务之间的界限也在不断变化。Grok Bot 反映了我们认为这条界限当前所处的位置。从最初的探索到最终发布,设计 Grok Bot 的过程就是寻找这条界限,并让界面随之演进的过程。随着智能体(Agent)承担更多责任,界面对人的要求应该更少。

Designing Grok Bot for a world of persistent agents

为持久化智能体的世界设计 Grok Bot

How we designed Grok Bot for agents that persist beyond a single session — from a chat history to a Bot roster, presence, a computer of the Bot’s own, and work that starts without a prompt.

我们如何为超越单一会话持久存在的智能体设计 Grok Bot——从聊天记录到 Bot 名册、在线状态、Bot 自己的计算机,以及无需提示词即可开始的工作。

When we started designing Grok Bot, one of the central questions was how the interface should shape the relationship between user and agent. Most AI interfaces are organized around a chat session the user operates. Each session begins with setup, unfolds as the user looks on, and ends when the conversation stops.

当我们开始设计 Grok Bot 时,核心问题之一是界面应如何塑造用户与智能体之间的关系。大多数 AI 界面都围绕用户操作的聊天会话组织。每个会话以设置开始,在用户的注视下展开,并在对话停止时结束。

We wanted to design for an agent that persists beyond any one session and can carry responsibility on its own. That meant reconsidering some of the basic objects and signals of the interface, including what belongs in the sidebar, how an agent shows progress, and when its work should become visible.

我们想要设计一种超越任何单一会话持久存在并能独立承担责任的智能体。这意味着需要重新考虑界面的一些基本对象和信号,包括侧边栏应该包含什么、智能体如何展示进度,以及它的工作何时应该变得可见。

Search

搜索

Kenny

Kenny

7:34 PM

7:34 PM

Need your yes on the Friday all-hands deck.

需要你对周五全员大会演示文稿的确认。

Justin

Justin

8 intros drafted — sitting in the CRM till you send.

起草了 8 份介绍信——放在 CRM 里等你发送。

Luke

Luke

7:34 PM

7:34 PM

Inboxs at 3. Two need a reply today.

收件箱有 3 封邮件。今天有两封需要回复。

Website launch

网站发布

11:18 AM

11:18 AM

John: checkouts clean on staging, 3 bugs closed.

John:预发环境结账流程正常,已关闭 3 个 bug。

John

John

Yesterday

昨天

Reprod the checkout crash. Write-ups in the ticket.

复现了结账崩溃问题。详情已写在工单里。

Keith

Keith

Acmes wobbling. Drafted a Thursday check-in.

Acme 情况不稳定。起草了一份周四汇报。

Tyler

Tyler

9:04 AM

9:04 AM

14 receipts in. Still missing your Uber from Tuesday.

已录入 14 张收据。仍缺你周二的 Uber 记录。

Manuel

Manuel

2:20 PM

2:20 PM

Launch post is live. First 200 impressions.

发布文章已上线。前 200 次展示。

Jenny

Jenny

Tuesday

周二

3 places in SoMa. The Folsom 2bed is the one.

SoMa 区有 3 处房源。Folsom 的两居室就是那个。

Chang

Chang

10:12 AM

10:12 AM

Sourced 3. Skipped one already in your ATS.

筛选了 3 份。跳过了一份已在你的 ATS 中的。

Plugins

插件

Peng Zheng

Peng Zheng

Kenny

Kenny

9:41 AM

9:41 AM

Hey, what were the main takeaways from today’s design sync?

嘿,今天设计同步会的主要结论是什么?

The team spent most of the session reviewing the new onboarding flow.

团队花了大部分时间审查新的引导流程。

The biggest discussion was around the empty states. Sarah felt the current illustrations don’t match the new brand direction, and most of the room agreed. There was also a long back and forth about whether the progress indicator belongs in the header or the sidebar.

最大的讨论围绕空状态展开。Sarah 觉得当前的插画不符合新的品牌方向,屋里大多数人都同意。关于进度指示器应该放在页头还是侧边栏,大家也进行了长时间的反复讨论。

The overall mood was positive though. Most people think the flow is close to ready for a wider review.

不过整体情绪很积极。大多数人认为该流程已接近准备好进行更广泛的审查。

Did anyone bring up the Q3 roadmap?

有人提起第三季度路线图吗?

Yes, it came up twice. Marcus said the roadmap review is now expected in the first week of August, and Priya asked whether the onboarding work would land before or after that. No firm dates were settled in the meeting.

是的,提了两次。Marcus 说路线图审查现在预计在八月第一周进行,Priya 问引导工作会在那之前还是之后完成。会上没有确定确切日期。

Message Kenny

给 Kenny 发消息

Kenny’s screen

Kenny 的屏幕

Routines

例行程序

Morning briefing

早间简报

Every day at 8:00 AM

每天上午 8:00

Inbox cleanup

收件箱清理

Weekdays at 6:00 PM

工作日下午 6:00

Weekly team update

每周团队更新

Paused

已暂停

AI products have accumulated a large vocabulary in a short time. Chats, sessions, models, context windows, memories, system prompts, projects, skills, connectors, agents, tools, sandboxes, permissions, and automations all describe real parts of these systems.

AI 产品在短时间内积累了大量词汇。聊天、会话、模型、上下文窗口、记忆、系统提示词、项目、技能、连接器、智能体、工具、沙盒、权限和自动化都描述了这些系统的真实组成部分。

But exposing each one as a separate product concept asks users to understand more than they need to. We started by asking which concepts a person actually needs in order to work with an agent.

但是,将每一个都作为单独的产品概念暴露出来,会让用户去理解超出他们需求的内容。我们首先提出的问题是:一个人实际上需要哪些概念才能与智能体协作。

We kept coming back to five:

我们不断回到这五个概念:

  • Bots are persistent agents with their own identity, memory, runtime, and tools.
  • Bot 是持久化的智能体,拥有自己的身份、记忆、运行时和工具。
  • Chats are the conversational interface for working with a Bot.
  • 聊天是与 Bot 协作的对话界面。
  • Prompts give a Bot context or instructions. They can be used once, saved as Skills, or triggered automatically as Routines.
  • 提示词为 Bot 提供上下文或指令。它们可以被使用一次、保存为技能,或作为例行程序自动触发。
  • Tools let Bots access information and take action through software, APIs, connectors, the shell, or computer use.
  • 工具让 Bot 能够通过软件、API、连接器、命令行或计算机使用来访问信息并采取行动。
  • Artifacts are the documents, designs, code, data, and other durable outputs that Bots create or modify.
  • 产物是 Bot 创建或修改的文档、设计、代码、数据和其他持久化输出。

Everything else could remain beneath the interface until the user had a reason to care about it. The next question was which of these five objects should organize the product.

其他一切都可以隐藏在界面之下,直到用户有理由关心它们。下一个问题是,这五个对象中哪一个应该用来组织产品。

Chats are disposable. We start a conversation to solve a problem. It gets pushed down the sidebar. A week later, we start another one. You rarely go back beyond the most recent five.

聊天是一次性的。我们开始一段对话是为了解决一个问题。它被推到侧边栏下方。一周后,我们又开始了一段新的对话。你很少会翻看最近五条以外的记录。

That behavior is perfectly reasonable when the unit of interaction is a question. It becomes strange when the thing on the other side of the interaction is supposed to know you, remember previous work, and take responsibility over time.

当交互的单位是一个问题时,这种行为是完全合理的。但当交互的另一端应该了解你、记住之前的工作并随着时间的推移承担责任时,这就变得奇怪了。

So the main objects in Grok Bot are Bots, not conversations. A Bot has a name. It has an avatar and a title. It remembers its conversations with you. It has its own computer and tools. When you come back tomorrow, you are coming back to the same Bot.

因此,Grok Bot 中的主要对象是 Bot,而不是对话。Bot 有名字。它有头像和标题。它记得与你的对话。它有自己的计算机和工具。当你明天回来时,你回到的是同一个 Bot。

Project Acme

Project Acme

Draft a follow-up to Acme after Friday’s call

在周五通话后给 Acme 起草一封跟进邮件

Rewrite the pricing one-pager

重写定价单页

What should I ask in the security review?

我在安全审查中应该问什么?

Build a champion map from my call notes

根据我的通话记录构建支持者地图

Practice the demo with hard objections

用尖锐的反对意见练习演示

Compare these three competitor decks for me

帮我比较这三个竞争对手的演示文稿

Show more

显示更多

Kenny

Kenny

7:34 PM

7:34 PM

Need your yes on the Friday all-hands deck.

需要你对周五全员大会演示文稿的确认。

Justin

Justin

8 intros drafted — sitting in the CRM till you send.

起草了 8 份介绍信——放在 CRM 里等你发送。

John

John

Yesterday

昨天

Reprod the checkout crash. Write-ups in the ticket.

复现了结账崩溃问题。详情已写在工单里。

Keith

Keith

Acmes wobbling. Drafted a Thursday check-in.

Acme 情况不稳定。起草了一份周四汇报。

Once a Bot was something you maintain over time rather than a session you start, the way Bots appear in the product had to answer three questions at once:

一旦 Bot 成为你随着时间维护的东西,而不是你开启的会话,Bot 在产品中出现的方式就必须同时回答三个问题:

  • Who is this?
  • 这是谁?
  • What are they doing?
  • 他们在做什么?
  • How much do I need to know?
  • 我需要了解多少?

A roster only works if it can be scanned quickly. As the roster grows, we did not want people to have to read every name each time they opened the product. They should be able to recognize a Bot from its avatar almost peripherally.

名册只有在能被快速扫视时才起作用。随着名册增长,我们不希望人们每次打开产品时都要阅读每个名字。他们应该几乎在余光中就能从头像认出一个 Bot。

Bot avatar visual style explorations by Kenny Kuh and Peng Zheng

Bot 头像视觉风格探索,作者:Kenny Kuh 和 Peng Zheng

At the same time, we wanted to keep the avatars consistent enough to read as one system. We studied character systems across illustration, animation, games, and interface design, exploring everything from initials and emojis to pixel art, watercolor, claymorphism, Noritake-style line art, silhouettes, and identicons.

同时,我们希望保持头像足够一致,使其看起来像一个系统。我们研究了插画、动画、游戏和界面设计中的角色系统,探索了从首字母和表情符号到像素艺术、水彩、黏土风、Noritake 风格线条画、剪影和身份图标等各种风格。

Most approaches solved one side of the problem better than the other. Watercolor and clay gave individual Bots plenty of character but carried too much detail at sidebar scale. Simpler systems sat more naturally in the interface, but often left the Bots looking interchangeable.

大多数方法只能更好地解决问题的一面。水彩和黏土风格赋予了单个 Bot 丰富的个性,但在侧边栏的尺寸下细节过多。更简单的系统在界面中显得更自然,但往往让 Bot 看起来可以互换。

The system we landed on keeps the basic construction consistent, using simple shapes and expressive eyes, then introduces distinction through controlled variations and accessories. Each Bot remains recognizable at a glance without appearing to come from a different visual world.

我们最终确定的系统保持了基本结构的一致性,使用简单的形状和富有表现力的眼睛,然后通过受控的变化和配饰引入差异性。每个 Bot 都能一眼认出,而不会显得来自不同的视觉世界。

Once the avatar became the Bot’s identity, it was also the natural place to show state. A Bot may be idle, thinking, working, waiting, blocked, or done. We could have represented each state with a separate indicator, but that would have added another layer of UI for the user to interpret.

一旦头像成为 Bot 的身份标识,它自然也就成了展示状态的理想位置。一个 Bot 可能处于空闲、思考、工作、等待、阻塞或完成状态。我们本可以用单独的指示器来表示每种状态,但这会给用户增加一层需要解读的 UI。

Instead, we explored how much of the lifecycle the avatar itself could carry.

相反,我们探索了头像本身能够承载多少生命周期。

At rest, the Bot is calm and slightly curious. When work arrives, it acknowledges the task. As work begins, it kicks into gear. Its motion changes again when it is waiting or needs help, then settles once the work is done. The avatar now shows what the Bot is doing as well as which Bot it is.

休息时,Bot 平静且略带好奇。当任务到来时,它会确认任务。工作开始时,它便进入状态。当它在等待或需要帮助时,动作会再次改变,一旦工作完成就会平静下来。现在,头像不仅展示了这是哪个 Bot,还展示了 Bot 正在做什么。

IdleWorkingWaitingBlockedThinkingDone

空闲工作等待阻塞思考完成

Avatar motion system by Benji Taylor

头像动作系统,作者:Benji Taylor

A related design question was how much of the Bot’s execution to show. One approach would have been the standard “three animated dots” but that would have been too little information, making it hard for users to tell whether the Bot was working or stuck.

一个相关的设计问题是,应该展示多少 Bot 的执行过程。一种方法是使用标准的“三个动画点”,但这提供的信息太少了,使得用户很难判断 Bot 是在工作还是卡住了。

Searching the web

正在搜索网络

  • Environment ready 387ms
  • 环境就绪 387ms
  • Edited math.ts +14 −10
  • 编辑了 math.ts +14 −10
  • Ran focused tests npm test
  • 运行了聚焦测试 npm test
  • Ran type-check npm run
  • 运行了类型检查 npm run
  • Thought briefly
  • 简短思考
  • Searched code “toFixed”
  • 搜索了代码 “toFixed”
  • Read AGENTS.md
  • 读取了 AGENTS.md
  • Edited math.test.ts +6 −2
  • 编辑了 math.test.ts +6 −2
  • Ran full suite 212 passed
  • 运行了全套测试 212 个通过
  • Committed fix: clamp NaN
  • 提交了修复:clamp NaN
  • Environment ready 387ms
  • 环境就绪 387ms
  • Edited math.ts +14 −10
  • 编辑了 math.ts +14 −10
  • Ran focused tests npm test
  • 运行了聚焦测试 npm test
  • Ran type-check npm run
  • 运行了类型检查 npm run
  • Thought briefly
  • 简短思考
  • Searched code “toFixed”
  • 搜索了代码 “toFixed”
  • Read AGENTS.md
  • 读取了 AGENTS.md
  • Edited math.test.ts +6 −2
  • 编辑了 math.test.ts +6 −2
  • Ran full suite 212 passed
  • 运行了全套测试 212 个通过
  • Committed fix: clamp NaN
  • 提交了修复:clamp NaN

We also tried showing a short written description of the Bot’s current action, but once people could see one step, they wanted to see the rest. User research showed us that they were asking for that detail mainly for reassurance that the Bot was still working and on the right track.

我们也尝试过展示 Bot 当前动作的简短文字描述,但一旦人们能看到一个步骤,他们就想看到其余的步骤。用户研究表明,他们要求提供这些细节,主要是为了确信 Bot 仍在工作且方向正确。

In the final design, the avatar’s motion provides the first bit of reassurance by showing that the Bot is active. If someone wants to check what it is doing, they can hover to see its current action.

在最终设计中,头像的动作通过展示 Bot 处于活跃状态,提供了第一重安心感。如果有人想检查它在做什么,可以将鼠标悬停在上面查看其当前动作。

Each Bot has its own computer, which it can use to browse the web, work with files, and run software. This created another interface problem. How visible should that computer be and when should the user be able to control it?

每个 Bot 都有自己的电脑,它可以用它来浏览网页、处理文件和运行软件。这带来了另一个界面问题。这台电脑应该有多显眼,用户什么时候应该能够控制它?

We explored four arrangements:

我们探索了四种布局:

  • Floating window: kept the computer easy to reach but covered the conversation.
  • 浮动窗口:保持电脑易于访问,但遮挡了对话。
  • Side by side: made the work continuously visible and encouraged users to watch it.
  • 并排显示:使工作过程持续可见,并鼓励用户观看。
  • Modal: made checking in easy but treated the Bot’s workspace as a temporary interruption.
  • 模态窗口:使查看变得容易,但将 Bot 的工作区视为临时中断。
  • Full screen: gave the computer plenty of room but displaced the conversation entirely.
  • 全屏:给电脑提供了充足的空间,但完全取代了对话。

The more prominent we made the computer, the more the product encouraged users to supervise it. We decided it should remain the Bot’s workspace, with the interface providing different levels of access as the user needed them.

我们让电脑越显眼,产品就越鼓励用户去监督它。我们决定它应该保持为 Bot 的工作区,界面根据用户的需求提供不同级别的访问权限。

The final design has three levels, which allow the user to enter the Bot’s workspace without being drawn into operating it:

最终设计有三个级别,允许用户进入 Bot 的工作区,而不必被卷入操作中:

  • Status: the title-bar icon turns purple while the computer is active.
  • 状态:当电脑处于活动状态时,标题栏图标变为紫色。
  • Preview: opening it reveals a pinned side panel where the user can follow the work without leaving the conversation.
  • 预览:打开它会显示一个固定的侧边面板,用户可以在不离开对话的情况下跟进工作。
  • Takeover: when the Bot needs help, the user can open the computer full screen, take control, and then hand it back.
  • 接管:当 Bot 需要帮助时,用户可以全屏打开电脑,接管控制权,然后再交还给它。

We also designed wallpapers that shift throughout the day, becoming lighter in the morning and darker at night. The detail gives the Bot’s computer its own sense of time and makes it feel separate from the user’s desktop.

我们还设计了随时间变化的壁纸,在早晨变亮,在夜晚变暗。这个细节赋予了 Bot 的电脑自己的时间感,使其感觉上独立于用户的桌面。

Dynamic wallpaper by Kenny Kuh and Luke Barker.

动态壁纸由 Kenny Kuh 和 Luke Barker 创作。

It is closer to working with a coworker than operating a remote machine. You can tell that they are working, glance at their screen when you need context, and sit down when something requires your help.

这更像是与一位同事共事,而不是操作一台远程机器。你能看出他们正在工作,在需要上下文(Context)时可以瞥一眼他们的屏幕,并在有事情需要你帮忙时坐下来处理。

Early versions of Grok Bot responded to almost every request with prose. It described a five-day forecast instead of showing one and narrated a set of tasks instead of laying them out as a board. The user then had to restructure the answer. This led us to treat the form of a response as part of the answer.

Grok Bot 的早期版本几乎对每个请求都用散文(Prose)来回复。它会描述未来五天的天气预报,而不是直接展示出来;它会口述一组任务,而不是将它们作为看板(Board)排列出来。然后用户不得不重新组织这些答案。这促使我们将回复的形式视为答案的一部分。

To support this, we built inline cards and widgets into Grok Bot. A Bot can answer in prose when prose fits the information and use structured UI when it does not.

为了支持这一点,我们在 Grok Bot 中内置了内联卡片(Inline cards)和小部件(Widgets)。当散文适合传达信息时,Bot 可以用散文回复;当不适合时,则使用结构化 UI(Structured UI)。

New email

新邮件

Ready to send

准备发送

Frompeng@grokbot.app

发件人 peng@grokbot.app

Tosarah@acme.com

收件人 sarah@acme.com

SubjectMoving Friday’s design review to 2 PM

主题 将周五的设计评审改到下午 2 点

Hi Sarah,

你好 Sarah,

Could we move Friday’s design review from 11 AM to 2 PM? A client call came up and I don’t want to rush our discussion.

我们能把周五的设计评审从上午 11 点改到下午 2 点吗?刚接到一个客户电话,我不想仓促进行我们的讨论。

Thanks, Peng

谢谢, Peng

Send emailDiscard

发送邮件 丢弃

Inline chat widgets by Peng Zheng

内联聊天小部件,由 Peng Zheng 提供

The same principle applies to actions. When a Bot creates a Routine, changes a setting, or messages another Bot, the event can appear directly in the transcript. The user can open it when there is more to inspect.

同样的原则也适用于操作(Actions)。当 Bot 创建一个例行程序(Routine)、更改设置或向另一个 Bot 发送消息时,该事件可以直接显示在对话记录(Transcript)中。如果有更多需要查看的内容,用户可以将其展开。

9:41 AM

上午 9:41

Morning! Can you check in with everyone for me?

早上好!你能帮我跟所有人确认一下进度吗?

On it — pinging the team for status now

正在处理——现在就向团队询问状态

6 messages withKennyTylerandJenny

与 Kenny、Tyler 和 Jenny 的 6 条消息

All on track: Kenny shipped the landing page, Tyler sent this months invoices, and Jenny booked next weeks interviews. No blockers.

一切顺利:Kenny 发布了落地页,Tyler 发送了本月的发票,Jenny 预订了下周的面试。没有阻碍。

Love it, can you do this every morning?

太棒了,你每天早上都能这样做吗?

Created RoutineMorning Briefing

已创建例行程序 晨间简报

Done, your Morning Briefing will be here at 9:00 every day

完成,你的晨间简报每天早上 9:00 会发到这里

The result is a heterogeneous transcript in which conversation, system events, interactive objects, and visualizations share one timeline.

结果是一个异构的对话记录(Heterogeneous transcript),其中对话、系统事件、交互对象和可视化内容共享同一个时间线。

Once people create several Bots, the product also has to organize how those Bots work together. We needed to decide which context should belong to each role, how Bots should share context when their work overlaps, and how to coordinate them without turning the user into a dispatcher.

一旦人们创建了多个 Bot,产品就必须组织好这些 Bot 如何协同工作。我们需要决定哪些上下文(Context)应该属于每个角色,当 Bot 的工作重叠时它们应该如何共享上下文,以及如何协调它们而不把用户变成调度员(Dispatcher)。

We saw one answer emerge as people created more Bots. Some made a Chief of Staff Bot responsible for coordinating several specialists. They could give direction to one Bot instead of checking each one and routing every task themselves.

随着人们创建的 Bot 越来越多,我们看到了一种答案。有些人创建了一个“幕僚长”Bot(Chief of Staff Bot)来负责协调几位专家。他们可以向一个 Bot 下达指令,而不必逐个检查并亲自分配每项任务。

Giving Bots distinct roles also forced us to decide what each role should know. A legal Bot may need the history of an ongoing dispute, while a finance Bot may need years of financial records. Combining those histories into one large memory would make it harder to give each Bot the information relevant to its work.

赋予 Bot 不同的角色也迫使我们决定每个角色应该知道什么。一个法务 Bot 可能需要了解正在进行的纠纷的历史,而一个财务 Bot 可能需要数年的财务记录。将这些历史合并到一个庞大的记忆(Memory)中,会使向每个 Bot 提供与其工作相关的信息变得更加困难。

Capabilities and context therefore follow different boundaries in Grok Bot. Tools and Skills live at the account level because many Bots may need to browse the web, work with documents, or send email. Memory and Routines belong to the Bot because they reflect what that particular role knows and does over time. Put another way, capabilities can be shared broadly while context remains with the role that needs it.

因此,在 Grok Bot 中,能力(Capabilities)和上下文遵循不同的边界。工具(Tools)和技能(Skills)存在于账户级别,因为许多 Bot 可能都需要浏览网页、处理文档或发送电子邮件。记忆和例行程序(Routines)则属于特定的 Bot,因为它们反映了该特定角色随着时间推移所掌握的知识和所做的事情。换言之,能力可以被广泛共享,而上下文则保留在需要它的角色身边。

Some work crosses those role boundaries. Group chats provide shared context for a project or team while allowing each Bot to retain its specialized memory. A designer, engineer, PM, and data scientist can work in the same conversation, hand work to one another, and share what the project requires.

有些工作会跨越这些角色边界。群聊为项目或团队提供了共享的上下文,同时允许每个 Bot 保留其专门的记忆。设计师、工程师、项目经理(PM)和数据科学家可以在同一个对话中工作,相互交接工作,并共享项目所需的内容。

We considered adding dashboards, assignment boards, and explicit handoff controls to manage these groups. Each one gave the user more coordination work. Instead, coordinating Bots handle routine routing and bring the user in when a decision requires judgment.

我们曾考虑添加仪表盘(Dashboards)、任务分配看板(Assignment boards)和明确的交接控制(Handoff controls)来管理这些群组。但每一种都给用户增加了更多的协调工作。相反

到了项目后期,许多设计工作都涉及做减法。我们移除了窗口和面板控件、计算机视图选项,以及智能体(Agent)元数据。我们还设定了实际限制:每个账户大约 50 个 Bot,每个群聊 6 个 Bot。每一个决策最终都归结为同一个问题:这究竟是帮助人们更好地委派任务,还是又给他们增加了一项需要管理的事务?

Most agent sessions begin when a user sends a prompt. That leaves even a persistent Bot waiting for someone to activate it. Routines let users give a Bot a standing responsibility that runs on a schedule or in response to an event, such as watching an industry or preparing a briefing every morning. The user defines the work once, and the Routine activates the Bot when it needs to happen.

随着模型的不断进步,操作 AI 与向同事委派任务之间的界限也在不断变化。Grok Bot 反映了我们认为这条界限当前所处的位置。从最初的探索到最终发布,设计 Grok Bot 的过程就是寻找这条界限,并让界面随之演进的过程。随着智能体(Agent)承担更多责任,界面对人的要求应该更少。

We initially treated Routines as secondary configuration. As they became more important to autonomous work, we moved them into the Bot’s main interface. The transcript shows what ran and gives the user a place to review the result or handle an exception.

Every day at 8:00 AM

On weekdays at 8:00 AM

Every Monday at 9:00 AM

Monthly on the 1st at 8:00 AM

Every 30 minutes

On weekdays · 9:00 AM

Issue created in grokbot

Issue any event in all projects

Incident triggered on grokbot

Incident any event on grokbot

PR opened in grokbot

PR merged in spacexai

PR closed in acme

New push to main

Checks fail on PRs in grokbot

Label bug added in grokbot

Comment containing /fix in ops

Review approved in grokbot

Thread resolved in grokbot

Workflow deploy.yml fails

When webhook receives a POST

New messages in #grokbot

Reaction :eyes: added in #ops

Channel created matching dev

New messages in #design

Issue created in Roadmap

Issue status → In Review in Core

At end of cycle for Team Core

Issue status → Done in Design

New messages in Design team

Channel created in Design team

This also changes the role of conversation. A prompt can start a session, but so can a schedule, an event, or another Bot. Over time, more work may begin without the user being present at all.

By the end of the project, much of the design work involved taking things away. We removed window and panel controls, computer-view options, and agent metadata. We also set practical limits of roughly 50 Bots per account and six per group chat. Each decision came back to the same question: Did this help someone delegate, or did it give them one more thing to manage?

The line between operating an AI and delegating to a coworker keeps moving as models improve. Grok Bot reflects where we think it sits today. Designing Grok Bot from its earliest explorations through launch has been about finding that line and helping the interface change with it. As agents take on more responsibility, the interface should ask less of the person.

Designing Grok Bot for a world of persistent agents

How we designed Grok Bot for agents that persist beyond a single session — from a chat history to a Bot roster, presence, a computer of the Bot’s own, and work that starts without a prompt.

When we started designing Grok Bot, one of the central questions was how the interface should shape the relationship between user and agent. Most AI interfaces are organized around a chat session the user operates. Each session begins with setup, unfolds as the user looks on, and ends when the conversation stops.

We wanted to design for an agent that persists beyond any one session and can carry responsibility on its own. That meant reconsidering some of the basic objects and signals of the interface, including what belongs in the sidebar, how an agent shows progress, and when its work should become visible.

Search

Kenny

7:34 PM

Need your yes on the Friday all-hands deck.

Justin

8 intros drafted — sitting in the CRM till you send.

Luke

7:34 PM

Inboxs at 3. Two need a reply today.

Website launch

11:18 AM

John: checkouts clean on staging, 3 bugs closed.

John

Yesterday

Reprod the checkout crash. Write-ups in the ticket.

Keith

Acmes wobbling. Drafted a Thursday check-in.

Tyler

9:04 AM

14 receipts in. Still missing your Uber from Tuesday.

Manuel

2:20 PM

Launch post is live. First 200 impressions.

Jenny

Tuesday

3 places in SoMa. The Folsom 2bed is the one.

Chang

10:12 AM

Sourced 3. Skipped one already in your ATS.

Plugins

Peng Zheng

Kenny

9:41 AM

Hey, what were the main takeaways from today’s design sync?

The team spent most of the session reviewing the new onboarding flow.

The biggest discussion was around the empty states. Sarah felt the current illustrations don’t match the new brand direction, and most of the room agreed. There was also a long back and forth about whether the progress indicator belongs in the header or the sidebar.

The overall mood was positive though. Most people think the flow is close to ready for a wider review.

Did anyone bring up the Q3 roadmap?

Yes, it came up twice. Marcus said the roadmap review is now expected in the first week of August, and Priya asked whether the onboarding work would land before or after that. No firm dates were settled in the meeting.

Message Kenny

Kenny’s screen

Routines

Morning briefing

Every day at 8:00 AM

Inbox cleanup

Weekdays at 6:00 PM

Weekly team update

Paused

Rethinking the primitives

AI products have accumulated a large vocabulary in a short time. Chats, sessions, models, context windows, memories, system prompts, projects, skills, connectors, agents, tools, sandboxes, permissions, and automations all describe real parts of these systems.

But exposing each one as a separate product concept asks users to understand more than they need to. We started by asking which concepts a person actually needs in order to work with an agent.

We kept coming back to five:

  • Bots are persistent agents with their own identity, memory, runtime, and tools.

  • Chats are the conversational interface for working with a Bot.

  • Prompts give a Bot context or instructions. They can be used once, saved as Skills, or triggered automatically as Routines.

  • Tools let Bots access information and take action through software, APIs, connectors, the shell, or computer use.

  • Artifacts are the documents, designs, code, data, and other durable outputs that Bots create or modify.

Everything else could remain beneath the interface until the user had a reason to care about it. The next question was which of these five objects should organize the product.

From chat history to a Bot roster

Chats are disposable. We start a conversation to solve a problem. It gets pushed down the sidebar. A week later, we start another one. You rarely go back beyond the most recent five.

That behavior is perfectly reasonable when the unit of interaction is a question. It becomes strange when the thing on the other side of the interaction is supposed to know you, remember previous work, and take responsibility over time.

So the main objects in Grok Bot are Bots, not conversations. A Bot has a name. It has an avatar and a title. It remembers its conversations with you. It has its own computer and tools. When you come back tomorrow, you are coming back to the same Bot.

Project Acme

Draft a follow-up to Acme after Friday’s call

Rewrite the pricing one-pager

What should I ask in the security review?

Build a champion map from my call notes

Practice the demo with hard objections

Compare these three competitor decks for me

Show more

Kenny

7:34 PM

Need your yes on the Friday all-hands deck.

Justin

8 intros drafted — sitting in the CRM till you send.

John

Yesterday

Reprod the checkout crash. Write-ups in the ticket.

Keith

Acmes wobbling. Drafted a Thursday check-in.

Presence as interface

Once a Bot was something you maintain over time rather than a session you start, the way Bots appear in the product had to answer three questions at once:

  • Who is this?

  • What are they doing?

  • How much do I need to know?

Who is this

A roster only works if it can be scanned quickly. As the roster grows, we did not want people to have to read every name each time they opened the product. They should be able to recognize a Bot from its avatar almost peripherally.

Bot avatar visual style explorations by Kenny Kuh and Peng Zheng

At the same time, we wanted to keep the avatars consistent enough to read as one system. We studied character systems across illustration, animation, games, and interface design, exploring everything from initials and emojis to pixel art, watercolor, claymorphism, Noritake-style line art, silhouettes, and identicons.

Most approaches solved one side of the problem better than the other. Watercolor and clay gave individual Bots plenty of character but carried too much detail at sidebar scale. Simpler systems sat more naturally in the interface, but often left the Bots looking interchangeable.

The system we landed on keeps the basic construction consistent, using simple shapes and expressive eyes, then introduces distinction through controlled variations and accessories. Each Bot remains recognizable at a glance without appearing to come from a different visual world.

What are they doing

Once the avatar became the Bot’s identity, it was also the natural place to show state. A Bot may be idle, thinking, working, waiting, blocked, or done. We could have represented each state with a separate indicator, but that would have added another layer of UI for the user to interpret.

Instead, we explored how much of the lifecycle the avatar itself could carry.

At rest, the Bot is calm and slightly curious. When work arrives, it acknowledges the task. As work begins, it kicks into gear. Its motion changes again when it is waiting or needs help, then settles once the work is done. The avatar now shows what the Bot is doing as well as which Bot it is.

IdleWorkingWaitingBlockedThinkingDone

Avatar motion system by Benji Taylor

How much do I need to know

A related design question was how much of the Bot’s execution to show. One approach would have been the standard “three animated dots” but that would have been too little information, making it hard for users to tell whether the Bot was working or stuck.

Searching the web

  • Environment ready 387ms

  • Edited math.ts +14 −10

  • Ran focused tests npm test

  • Ran type-check npm run

  • Thought briefly

  • Searched code “toFixed”

  • Read AGENTS.md

  • Edited math.test.ts +6 −2

  • Ran full suite 212 passed

  • Committed fix: clamp NaN

  • Environment ready 387ms

  • Edited math.ts +14 −10

  • Ran focused tests npm test

  • Ran type-check npm run

  • Thought briefly

  • Searched code “toFixed”

  • Read AGENTS.md

  • Edited math.test.ts +6 −2

  • Ran full suite 212 passed

  • Committed fix: clamp NaN

We also tried showing a short written description of the Bot’s current action, but once people could see one step, they wanted to see the rest. User research showed us that they were asking for that detail mainly for reassurance that the Bot was still working and on the right track.

In the final design, the avatar’s motion provides the first bit of reassurance by showing that the Bot is active. If someone wants to check what it is doing, they can hover to see its current action.

Their computer, not yours

Each Bot has its own computer, which it can use to browse the web, work with files, and run software. This created another interface problem. How visible should that computer be and when should the user be able to control it?

We explored four arrangements:

  • Floating window: kept the computer easy to reach but covered the conversation.

  • Side by side: made the work continuously visible and encouraged users to watch it.

  • Modal: made checking in easy but treated the Bot’s workspace as a temporary interruption.

  • Full screen: gave the computer plenty of room but displaced the conversation entirely.

The more prominent we made the computer, the more the product encouraged users to supervise it. We decided it should remain the Bot’s workspace, with the interface providing different levels of access as the user needed them.

The final design has three levels, which allow the user to enter the Bot’s workspace without being drawn into operating it:

  • Status: the title-bar icon turns purple while the computer is active.

  • Preview: opening it reveals a pinned side panel where the user can follow the work without leaving the conversation.

  • Takeover: when the Bot needs help, the user can open the computer full screen, take control, and then hand it back.

We also designed wallpapers that shift throughout the day, becoming lighter in the morning and darker at night. The detail gives the Bot’s computer its own sense of time and makes it feel separate from the user’s desktop.

Dynamic wallpaper by Kenny Kuh and Luke Barker.

It is closer to working with a coworker than operating a remote machine. You can tell that they are working, glance at their screen when you need context, and sit down when something requires your help.

The shape of information

Early versions of Grok Bot responded to almost every request with prose. It described a five-day forecast instead of showing one and narrated a set of tasks instead of laying them out as a board. The user then had to restructure the answer. This led us to treat the form of a response as part of the answer.

To support this, we built inline cards and widgets into Grok Bot. A Bot can answer in prose when prose fits the information and use structured UI when it does not.

New email

Ready to send

Frompeng@grokbot.app

Tosarah@acme.com

SubjectMoving Friday’s design review to 2 PM

Hi Sarah,

Could we move Friday’s design review from 11 AM to 2 PM? A client call came up and I don’t want to rush our discussion.

Thanks, Peng

Send emailDiscard

Inline chat widgets by Peng Zheng

The same principle applies to actions. When a Bot creates a Routine, changes a setting, or messages another Bot, the event can appear directly in the transcript. The user can open it when there is more to inspect.

9:41 AM

Morning! Can you check in with everyone for me?

On it — pinging the team for status now

6 messages withKennyTylerandJenny

All on track: Kenny shipped the landing page, Tyler sent this months invoices, and Jenny booked next weeks interviews. No blockers.

Love it, can you do this every morning?

Created RoutineMorning Briefing

Done, your Morning Briefing will be here at 9:00 every day

The result is a heterogeneous transcript in which conversation, system events, interactive objects, and visualizations share one timeline.

Organizing intelligence

Once people create several Bots, the product also has to organize how those Bots work together. We needed to decide which context should belong to each role, how Bots should share context when their work overlaps, and how to coordinate them without turning the user into a dispatcher.

We saw one answer emerge as people created more Bots. Some made a Chief of Staff Bot responsible for coordinating several specialists. They could give direction to one Bot instead of checking each one and routing every task themselves.

Giving Bots distinct roles also forced us to decide what each role should know. A legal Bot may need the history of an ongoing dispute, while a finance Bot may need years of financial records. Combining those histories into one large memory would make it harder to give each Bot the information relevant to its work.

Capabilities and context therefore follow different boundaries in Grok Bot. Tools and Skills live at the account level because many Bots may need to browse the web, work with documents, or send email. Memory and Routines belong to the Bot because they reflect what that particular role knows and does over time. Put another way, capabilities can be shared broadly while context remains with the role that needs it.

Some work crosses those role boundaries. Group chats provide shared context for a project or team while allowing each Bot to retain its specialized memory. A designer, engineer, PM, and data scientist can work in the same conversation, hand work to one another, and share what the project requires.

We considered adding dashboards, assignment boards, and explicit handoff controls to manage these groups. Each one gave the user more coordination work. Instead, coordinating Bots handle routine routing and bring the user in when a decision requires judgment.

Work that keeps moving

Most agent sessions begin when a user sends a prompt. That leaves even a persistent Bot waiting for someone to activate it. Routines let users give a Bot a standing responsibility that runs on a schedule or in response to an event, such as watching an industry or preparing a briefing every morning. The user defines the work once, and the Routine activates the Bot when it needs to happen.

We initially treated Routines as secondary configuration. As they became more important to autonomous work, we moved them into the Bot’s main interface. The transcript shows what ran and gives the user a place to review the result or handle an exception.

Every day at 8:00 AM

On weekdays at 8:00 AM

Every Monday at 9:00 AM

Monthly on the 1st at 8:00 AM

Every 30 minutes

On weekdays · 9:00 AM

Issue created in grokbot

Issue any event in all projects

Incident triggered on grokbot

Incident any event on grokbot

PR opened in grokbot

PR merged in spacexai

PR closed in acme

New push to main

Checks fail on PRs in grokbot

Label bug added in grokbot

Comment containing /fix in ops

Review approved in grokbot

Thread resolved in grokbot

Workflow deploy.yml fails

When webhook receives a POST

New messages in #grokbot

Reaction :eyes: added in #ops

Channel created matching dev

New messages in #design

Issue created in Roadmap

Issue status → In Review in Core

At end of cycle for Team Core

Issue status → Done in Design

New messages in Design team

Channel created in Design team

This also changes the role of conversation. A prompt can start a session, but so can a schedule, an event, or another Bot. Over time, more work may begin without the user being present at all.

The disappearing interface

By the end of the project, much of the design work involved taking things away. We removed window and panel controls, computer-view options, and agent metadata. We also set practical limits of roughly 50 Bots per account and six per group chat. Each decision came back to the same question: Did this help someone delegate, or did it give them one more thing to manage?

The line between operating an AI and delegating to a coworker keeps moving as models improve. Grok Bot reflects where we think it sits today. Designing Grok Bot from its earliest explorations through launch has been about finding that line and helping the interface change with it. As agents take on more responsibility, the interface should ask less of the person.

📋 讨论归档

讨论进行中…