从连接到意图:互联网、AI 与人的下一站 / From Connection to Intent: The Internet, AI, and What Comes Next for Humans
从连接到意图:互联网、AI 与人的下一站 / From Connection to Intent: The Internet, AI, and What Comes Next for Humans
本文包含中文与英文两个版本。
This post includes both Chinese and English versions.
[CN] 中文版
思考的起点
最近在想,互联网到底是在不断变得更复杂,还是在不断变得更简单。
从文字到图片、视频,再到 AR 和 VR,看起来是媒介越来越丰富;从手动操作软件到使用 AI,看起来是工具越来越强;从门户到平台,再到个人 Agent,看起来又是权力结构在变化。
如果只看产品,会得到一串很长的名单。如果往底层看,这些变化大概可以归纳成三条同时发生的路线:
- 人如何感知世界;
- 人如何把意图变成结果;
- 谁拥有连接、数据和分配资源的权力。
这三条线并不是严格的先后关系。文字时代并没有结束,平台也没有消失,AI 也还没有真正接管现实。所谓“下一阶段”,更多是旧结构与新结构同时存在,并且彼此争夺主导权。
一、从文字到空间:媒介正在接近人的感官
文字是一种低带宽、高抽象的媒介。
我们看到文字以后,需要在脑海中重新构建一个世界。这个过程比较慢,却也保留了思考的距离。很多复杂的概念,只有经过这种二次加工,才真正变成自己的理解。
图片和视频提高了信息的带宽,也降低了理解的门槛。人不必先掌握太多抽象符号,内容就可以直接进入视野。代价是,内容也更容易绕过判断,直接占用注意力。
AR 和 VR 继续沿着这条路线向前走。它们不只是更大的屏幕,而是试图把信息放回空间:地图在街道上,提示出现在物体旁边,数字内容与现实环境叠在一起。
这条路线的终点,可能不是“看到更多信息”,而是“看不见界面”。人不再明显地进入互联网,互联网也不再表现为一个单独的页面。
但媒介越自然,注意力越需要被保护。技术可以降低获取信息的成本,却未必降低理解信息的成本。看见更多,不等于判断得更好。
二、从操作工具到表达意图:生产力开始自动驾驶
AI 之前,人和软件之间主要是操作关系。
我们学习菜单、按钮、快捷键和工作流,软件把复杂能力拆成一个个可以点击的动作。Word、Photoshop、表格软件和各种专业工具,本质上都要求人适应机器提供的界面。
AI 时代的变化,是人开始从“怎么做”转向“想得到什么”。
过去可能需要这样工作:打开文件,查找资料,整理表格,修改内容,再导出结果。以后更自然的表达可能是:把这批资料整理成一份可以交付的方案,并告诉我其中的风险。
中间的步骤由 Agent 完成。
这会改变生产力的门槛。熟练操作依然有价值,但更重要的能力会逐渐变成:
- 能不能把问题说清楚;
- 能不能把目标拆开;
- 能不能判断结果是否真的有效;
- 能不能为结果及其后果负责。
AI 不只是放大执行能力,也会放大目标本身的错误。一个没有想清楚问题的人,可能更快地得到一份看起来完整、实际上没有解决问题的答案。
所以,未来真正稀缺的可能不是“会不会使用某个模型”,而是能不能定义一个值得解决的问题。
三、从门户到平台,再到个人 Agent:权力可能重新移动
互联网早期是门户时代。少数入口负责编辑和分发,用户主要是信息的接受者。
后来进入平台时代。搜索、社交、电商和内容平台聚集了数据,也聚集了分配注意力的权力。我们拥有账号,却不一定真正拥有关系、内容和行为轨迹。
下一步可能是 Agent 时代。
如果每个人都有一个长期理解自己目标、习惯和边界的 Agent,那么人与互联网的关系就会发生变化。过去是人进入不同 App 寻找服务,未来可能是个人 Agent 去调用整个互联网。
这会带来一种新的个人主权,也会带来新的风险。一个 Agent 知道得越多,就越像个人的数字延伸;它的权限越大,错误和滥用造成的损失也越大。
真正的数字主权,不只是拥有一个自己的 Agent,还包括知道它记住了什么、代表谁做决定,以及什么情况下必须停下来询问人。
四、典范不是卖得最好,而是改变了常识
对话中有一个问题让我印象比较深:一个时代的典范究竟是什么。
我觉得,典范不只是卖得好的产品,也不只是最早出现的技术。它更像一种新的常识:让原本复杂、昂贵的能力变得普通,并且让人很难再回到过去的生活方式。
Google 降低了获取信息的成本,智能手机降低了连接网络的成本,短视频改变了内容分发的方式,ChatGPT 则降低了调用部分知识和创作能力的门槛。
它们共同的特点,是把某种能力从少数人的技能,变成多数人的日常动作。
因此,下一代典范也许不是一个功能更多的软件,而是一个让界面消失的系统:
- 不再逐个打开 App,而是直接表达意图;
- 不再一直盯着屏幕,而是在现实环境中获得辅助;
- 不再只提供信息,而是直接交付可以使用的结果。
五、大公司争夺的不是同一块市场,而是同一个入口
从这个角度看,英伟达、OpenAI、Meta、Apple 和 Google 的动作虽然不同,争夺的却可能是同一件事:下一代互联网的默认入口。
英伟达更接近底层。它不决定人想做什么,却决定多少计算能够被现实地运行。
OpenAI 更接近生产力入口。它试图把应用的操作藏在自然语言和 Agent 之后,让人从使用一堆工具,变成调用一组能力。
Meta 和 Apple 更接近感知入口。它们关注的不是普通屏幕,而是人如何看见现实、如何进入空间,以及下一代设备是否足够自然。
这些公司的成功并不是确定的。基础设施可能遇到成本压力,模型可能遇到隐私和开源竞争,硬件可能卡在重量、续航和社交接受度,平台也可能不愿意把入口让给别人的 Agent。
对个人来说,更现实的机会不在于复制它们的规模,而在于占据一个具体的意图:为某个行业构建垂直 Agent,为真实的空间和线下场景做内容,让自己的经验、审美和工作流变成个人 AI 资产,或者成为某个领域的可信来源。
六、失业首先意味着任务被替代,最后才是岗位被重新定义
关于失业,最容易出现两种极端判断:要么认为所有工作都会消失,要么认为 AI 只是普通工具,不会改变什么。
更可能发生的,是任务先被替代,岗位随后被重新组合。
基础翻译、摘要、绘图、代码和合规检查中的一部分工作,会更快地交给机器。但一个岗位通常不只是任务清单,它还包括沟通、判断、责任以及面对不确定性的能力。
真正的变化,是一个人可以独立完成的事情变多了。过去需要几个人配合的小项目,可能由一个人带着一组 Agent 完成。这会提高个人的上限,也会让组织结构变薄。
如果机器创造了更多财富,社会是否仍然只能通过工资分配资源?全民基本收入、数据红利、算力红利和机器人税,都可能成为制度讨论的一部分。但技术不会自动带来公平,谁拥有生产系统、谁决定机器创造的价值如何分配,仍然是核心问题。
即使基本生活可以被保障,人仍然需要一种“我有用”的感觉。未来需要重新分配的不只是收入,也包括参与感、尊严和社会价值。
七、以人为本不会消失,只是人从执行中心移向意图上游
过去估算一个公司的大小,通常从人口、需求和购买力出发:
市场空间 = 人口数量 × 人均需求 × 购买力
AI Agent 出现以后,产业的参与者可能不再只有人。后台可以运行大量数字代理,软件的规模也可能由任务次数、API 调用、数据流动和算力消耗决定。
但这并不意味着人被移出了公式。
机器可以执行更多任务,却不会凭空产生最初的欲望。它可以优化路径,却不能替我们决定什么值得追求。它可以比较所有方案,却不能替我们承担选择失败的后果。
未来可以先用一个不那么精确、但有助于思考的公式表示:
产业价值 = 人类的意图定义 × Agent 的执行规模 × 算力与资源效率
人的位置会从执行中心,逐渐向意图和责任的上游移动。
功能会越来越便宜,信任、真实体验和“人味”反而可能变得更贵。以人为本不会消失,只是从整个经济系统的默认尺度,变成最难被复制的一部分。
结语
互联网的下一阶段,可能不是让人看到更多、操作更快,而是让更多事情在没有明显界面的情况下被完成。
但工具越强,越需要有人回答最初的问题:我们为什么要做这件事?谁会从中受益?如果结果出错,谁来负责?
机器会不断扩大“能够完成什么”的边界。
而人真正需要守住的,也许正是“什么值得完成”的判断。
[EN] English Version
Where the Question Begins
Recently, I have been thinking about whether the Internet is becoming more complicated, or quietly becoming simpler.
From text to images and video, and then to AR and VR, media keeps becoming richer. From manual software operation to AI, tools keep becoming stronger. From portals to platforms and then to personal Agents, the structure of power is changing as well.
Looking underneath, these changes form three paths unfolding at the same time: how people perceive the world, how people turn intent into results, and who owns connection, data, and the power to distribute resources.
I. From Text to Space
Text is a low-bandwidth, high-abstraction medium. We reconstruct a world in our minds after reading. It is slow, but it leaves room for thought.
Images and video increase bandwidth and lower the barrier to understanding. AR and VR continue this direction by putting information back into space. The destination may not be seeing more information, but no longer seeing the interface.
The more natural the medium becomes, the more attention needs protection. Technology can lower the cost of accessing information without lowering the cost of understanding it.
II. From Operating Tools to Expressing Intent
Before AI, people learned menus, buttons, shortcuts, and workflows. The AI-era shift is from explaining how to do something to expressing what should be achieved.
The intermediate steps can be handled by an Agent. Defining a problem, decomposing a goal, judging an outcome, and taking responsibility for consequences may matter more than knowing a particular model.
AI amplifies execution, but it also amplifies flawed objectives. The scarce skill of the future may be defining a problem worth solving.
III. From Platforms to Personal Agents
The early Internet was the portal era. The platform era then accumulated data and the power to distribute attention. We owned accounts, but not necessarily our relationships, content, or behavioral histories.
If every person has an Agent that understands their goals, habits, and boundaries, a personal Agent may call upon the Internet on our behalf. Apps may remain, but move into the background as capabilities that Agents can invoke.
Digital sovereignty is not merely owning an Agent. It means knowing what it remembers, whom it represents, and when it must stop and ask a human.
IV. What Makes a Paradigm
A paradigm is not simply the best-selling product. It is a new common sense that makes a complex ability ordinary and makes the old way difficult to return to.
Google lowered the cost of finding information. Smartphones lowered the cost of staying connected. Short video changed distribution. ChatGPT lowered the barrier to accessing parts of knowledge and creative capability.
The next paradigm may be a system that makes the interface disappear: expressing intent instead of opening every app, receiving assistance in the physical world instead of staring at a screen, and receiving usable results instead of merely receiving information.
V. The Competition for the Next Entrance
NVIDIA, OpenAI, Meta, Apple, and Google operate in different markets, but they may be competing for the same position: the default entrance to the next Internet.
NVIDIA is closest to the foundation. OpenAI is closer to the productivity entrance. Meta and Apple are closer to the perception entrance. Google has the information and browser entrances, but must reconsider them if Agents deliver answers directly.
None of these outcomes is certain. Infrastructure faces cost pressure, models face privacy and open-source competition, hardware faces weight and social acceptance, and platforms may not want to surrender the entrance to someone else’s Agent.
For individuals, the realistic opportunity is to own a specific intent: build a vertical Agent, create for real offline situations, turn accumulated experience into a personal AI asset, or become a trusted source in a narrow field.
VI. Unemployment and Distribution
A more likely path is that tasks are replaced first, and jobs are recombined afterward. Parts of translation, summarization, drawing, coding, and compliance work will move to machines, while communication, judgment, responsibility, and uncertainty remain important.
One person may be able to complete much more alone. This raises individual productivity while making organizations thinner. It also weakens the old link between labor and output.
Universal basic income, data dividends, compute dividends, and robot taxes may become part of the institutional debate. Technology will not create fairness by itself. Ownership of the productive system and control over distribution remain the central questions.
Even if basic needs are secured, people will still need a sense of usefulness. The future may need to distribute not only income, but also participation, dignity, and social meaning.
VII. Humanity Moves Upstream
With AI Agents, market participants may no longer be only human beings. Digital agents may run in the background, while software scale is measured through tasks, API calls, data flows, and compute.
This does not remove humans from the formula. Machines can execute more tasks, but they do not invent the original desire. They can optimize a path, but they cannot decide what is worth pursuing or bear the consequences of a failed choice.
Industry value = human intent-definition × Agent execution scale × compute and resource efficiency
Humanity moves from the center of execution toward the upstream position of intent and responsibility. Function becomes cheaper. Trust, real experience, and human warmth may become more expensive.
Conclusion
The next stage of the Internet may be about completing more tasks without an obvious interface.
But the stronger the tool becomes, the more important the original questions are: Why are we doing this? Who benefits? Who is responsible when the result is wrong?
Machines will keep expanding the boundary of what can be completed.
What humans may need to protect is the judgment of what is worth completing.