NOEMATRIX LAB研究博客 NO.1Research Blog NO.1
Noe-0研究概览 Noe-0 Research Preview
全流程无本体数据,
突破灵巧性瓶颈。
End-to-end non-embodied data.
Breaking through the dexterity bottleneck.
目前机器人的后训练通常依赖遥操作,轨迹中的停顿、分段与反复修正可能压缩预训练阶段获得的动作知识。Noe-0 在预训练与后训练中统一使用无本体数据,使自然动作直接迁移为机器人任务能力。 Conventional post-training often relies on robot teleoperation. Pauses, segmentation, and repeated corrections in teleoperated trajectories can narrow the motion knowledge acquired during pre-training. Noe-0 uses non-embodied data in both stages, allowing natural motion knowledge to transfer more directly into robotic task capabilities.
引言Introduction
从零构建数据、模型与基础设施的全栈闭环。A full-stack loop built from scratch across data, models, and infrastructure.
具身基础模型,依赖全栈系统。Embodied foundation models demand a full-stack system.
具身通用基础模型不是单一的模型架构问题,而是一项全栈系统工程。真实世界的数据生产、物理世界与动作的联合建模,以及支撑规模化迭代的基础设施,都需要从零开始、协同建设,才能让数据、训练与反馈形成高效、完整的闭环。Building a general-purpose foundation model for embodied intelligence is not solely a model-architecture problem; it is a full-stack systems problem. Real-world data production, joint world–action modeling, and infrastructure for iteration at scale must all be built from scratch and developed together, so that data, training, and feedback form an efficient, complete loop.
经过过去数月的工作,我们很高兴在这篇研究概览中分享 Noe-0 的近期进展与当前验证结果。这里展示了已经完成的关键系统环节:无本体数据的采集与治理、世界—动作联合建模,以及连接数据生产与管理的基础设施。我们也附上了一些真实数据样例、实际的系统流程和训练曲线,共同呈现这条路径目前已经走到哪里。After several months of work, we are pleased to share Noe-0’s recent progress and the results validated to date. This Research Preview presents the key parts of the system now in place: non-embodied data collection and governance, joint world–action modeling, and infrastructure connecting data production with natural-language querying. Together, the real-world samples, system workflows, and training curves show how far this path has progressed.
数据
Data
真实世界采集、整理与版本化。
Capture, curation, and versioning in the real world.
模型
Model
联合学习世界变化与动作。
Jointly learn world dynamics and action.
基础设施
Infrastructure
让每次处理、训练与验证可追踪。
Keep processing, training, and validation traceable.
数据Data
从自然动作到可扩展数据。From natural motion to scalable data.
动作质量,始于采集。Motion quality begins with how it is captured.
人类在操作过程中,手部会连续调整姿态、接触力与节奏。RoboPocket 使用同步的头部与双腕视角,在不依赖机器人遥操作的前提下,记录这种自然、灵巧的操作过程。Human hands continuously adjust pose, contact, and rhythm. RoboPocket uses synchronized head-mounted and dual-wrist views to capture natural dexterous behavior without robot teleoperation.
RoboPocket / natural capture
保留动作,再形成能力。Preserve motion. Build capability.
三个同步视角同时记录手、物体与环境变化:头部视角保留任务意图,两路手腕视角补充接触与局部动作。Three synchronized views capture hands, objects, and scene change together: the head view preserves task intent while two wrist views reveal contact and local motion.
预训练从这些无本体数据中学习细粒度手部运动、物体状态变化与动作节奏;它看到的不是遥操作能够复现的机械性的运动轨迹,而是人类操作本身更自然的动作分布。During pre-training, the model learns fine-grained hand motion, object-state changes, and temporal dynamics from this non-embodied data. This exposes the model to a broader distribution of human manipulation than the mechanically constrained trajectories typically produced by teleoperation.
如果后训练重新换成停顿、机械性动作和频繁修正的遥操作轨迹,早期学到的动作知识就会被压缩到更窄的分布中。Noe-0 在两个阶段沿用无本体数据,让自然动作的灵巧性可以完整保留在模型中。If post-training switches back to hesitant, mechanically constrained, and frequently corrected teleoperation trajectories, earlier motion knowledge is compressed into a narrower distribution. Noe-0 uses non-embodied data during both stages to reduce the loss of natural motion characteristics during post-training.
采集自然动作Capture natural motion
头部与双腕同步记录任务意图、局部接触和连续动作。Head and dual-wrist views preserve intent, contact, and continuous motion.
突破遥操作瓶颈Break through the teleoperation bottleneck
后训练不再用机械化轨迹重新限制预训练知识。Post-training does not re-limit pre-training knowledge with mechanical trajectories.
保持动作语言连续Keep one action language
同一种无本体动作分布从预训练延续到任务能力形成。One non-embodied action distribution continues into task-level capability.
集中式数据采集受限于标准化搭建的采集场景,操作对象大多批量统一采购,因此难以自然获得足够的数据多样性。若将数据采集下沉至采集人员的生活空间,环境内日积月累的包装、工具、容器和日用品,可覆盖材质、形态及使用状态下的各类长尾样本。Centralized data collection is constrained by standardized collection environments, where objects are typically purchased in uniform batches, making sufficient diversity difficult to obtain naturally. Moving collection into collectors’ living spaces brings in accumulated packaging, tools, containers, and everyday items, covering long-tail samples across materials, forms, and conditions of use.
累计数据时长
Cumulative data hours
97,387.7 h
工具测量
Tool measurement

纸张整理
Paper handling

电子部件装配
Device assembly
02.C / 真实世界数据02.C / REAL-WORLD DATA
真实世界数据样例。Real-world data samples.
覆盖精细装配、工具使用、柔性物体、分类归置与多阶段操作。点击按钮可随机查看另一组数据样例。Spanning precision assembly, tool use, deformable objects, sorting, and multi-stage manipulation. Use the button to view another selection.
模型Model
同时进行隐式世界想象和工作规划。Perform implicit world imagination and task planning simultaneously.
一边隐式想象世界,
一边规划下一步动作。Imagine the world implicitly while planning the next
action.
Noe-0 接收图像与任务语义后,在同一个表示中同时进行隐式世界想象与动作规划,并输出动作。它并非“先想象、再规划”的串行流水线,而是在行动发生时持续保持对世界变化的预测。Given images and task semantics, Noe-0 performs implicit world imagination and action planning in one shared representation while producing actions. It is not a serial imagine-then-plan pipeline; prediction of world change remains active as actions unfold.
隐式想象为行动规划提供对象状态与因果上下文;行动信号又反过来约束模型,只关注真正影响任务进程的变化。两者共享时间步与表示空间,而非通过前后两个独立模块顺序传递。Implicit imagination supplies object state and causal context to action planning; action signals in turn constrain the model to changes that matter to task progress. The two share a time step and representation space rather than passing through separate modules in sequence.
执行动作后,新观察进入下一轮联合推理。完整闭环是“观察 → 隐式想象与行动规划同时发生 → 执行动作 → 再观察”,从而在长时序任务中持续更新世界状态与控制决策。After the action executes, a new observation enters the next joint inference step. The full loop is observe → imagine implicitly and plan simultaneously → execute → observe again, continuously updating world state and control decisions across long-horizon tasks.
01图像 + 任务语义
Images + task semantics
Noe-0
噪声动作 Token 可见噪声视频 Token;隐式世界想象与动作规划并行发生。
Noisy action tokens attend to noisy video tokens while world imagination and action planning run in parallel.
03动作输出 + 下一次观察
Action output + next observation
用模型扩展,
验证数据有效性。Use scaling behavior in model training to validate data effectiveness.
在固定的验证集设置下,10%、20% 与 100% 数据量实验呈现 loss 的有序下降。Under a fixed validation-set protocol, the 10%, 20%, and 100% data-volume runs show an ordered decline in loss.
固定验证集 / 越低越好FIXED VALIDATION SET / LOWER IS BETTER
完整评测轨迹Complete traces
从 56k 步开始,每 2k 步记录一次验证动作损失。Validation action loss is recorded every 2k steps from 56k onward, so endpoints never stand in for the process.
固定验证集Fixed validation set
三组实验使用同一固定验证集设置。All three runs use the same fixed validation set.
仅改变数据量Data volume only
三组实验仅改变训练数据量。Only training-data volume changes.
基础设施Infrastructure
让数据生产与模型迭代形成可追踪闭环。A traceable loop between data production and model iteration.
从端侧采集,
到可交付的训练资产。From edge capture to training-ready assets.
端侧数采系统 RoboPocket 的生产链把端侧三视角采集、任务管理、采集工具、数据管理平台 DM3 云端入库、质量检查、人工审核、标注审核与数据交付连成一条可追踪链路。The edge data-collection system RoboPocket connects three-view edge capture, task management, capture tooling, cloud ingest into the data management platform DM3, quality checks, human and annotation review, and data delivery in one traceable path.
端侧负责把同步、完整的多视角任务数据带回云端;云端连续完成入库、自动检查与人工判断。每一次质量结论都会与任务、采集片段和数据版本保持关联,便于回溯与修正。The edge returns synchronized, complete multi-view task data; the cloud performs ingest, automated checks, and human review. Every quality decision stays linked to its task, capture, and data version for traceability and revision.
在已审批数据切片的同口径对照中,半自动标注用时较纯人工减少约 22%;自动化处理高频判断,人处理难例;交付后的训练资产继续与配置、检查点和评测反馈关联,把模型中暴露的问题送回下一轮数据工作。In a like-for-like comparison on an approved data slice, semi-automated annotation reduced time by approximately 22%; Automation handles recurring decisions while people resolve hard cases. Delivered assets remain linked to training configurations, checkpoints, and evaluation feedback, returning model failures to the next data cycle.
端侧采集
Edge capture
头部 + 双腕同步
head + dual-wrist sync
任务管理
Task management
采集工具
Capture tool
RoboPocket
云端入库
Cloud ingest
DM3
云端质检
Cloud QC
人工审核
Human review
标注审核
Annotation review
AI 预标注 + 人工优化
AI pre-labeling + human refinement
数据交付
Data delivery
语义聚类 + 数据筛选
clustering + filtering
Data Agent
以自然语言连接数采管理员与数采员。
A natural-language bridge between data collection managers and data collectors.
04.B / Data Agent
把分散的数据操作,
组织成一条可追踪工作流。Bring distributed data operations into one traceable workflow.
Data Agent 连接任务生成、数据查询与统计、质量问题处理、消息通知和权限管理。用户可以直接用自然语言查询采集进度、任务状态与质量问题。它不替代人的质量判断,而是把分散的操作组织成同一条可回溯工作流,让每次修正直接服务于后续数据版本。The Data Agent connects task generation, query and statistics, quality issue resolution, notifications, and access control. Users can query collection progress, task status, and quality issues directly in natural language. It does not replace human quality judgment; it organizes scattered operations into a traceable workflow so every correction reaches later data versions.
让重复判断交给半自动链路Make recurring judgment semi-automatic
任务理解、质量检查、预标注与置信度判断形成统一入口,让重复判断进入稳定的半自动链路。Task interpretation, quality review, pre-annotation, and confidence estimation share one entry point, moving recurring decisions into a stable semi-automated workflow.
让训练资产可追踪、可复用Keep training assets traceable and reusable
数据版本、训练配置、检查点与评测结果保持关联,使每次迭代都有完整依据。Data versions, training configurations, checkpoints, and evaluations remain linked so every iteration has a complete evidence trail.
让每轮训练改善下一轮数据Let every model improve the next dataset
模型反馈驱动任务补齐与采集修正,使数据生产与能力增长形成闭环。Model feedback drives task coverage and capture correction, closing the loop between data production and capability growth.
展望Outlook
闭环不是终点,而是更快逼近通用性的起点。The loop is not the finish line; it is the engine for approaching generality faster.
闭环迭代,持续突破能力边界。Close-loop iteration, continuously pushing the boundaries of capability.
目前,数据生产、世界—动作联合建模与反馈基础设施已经连接成一条可工作的链路。至于更开放的语言跟随、跨场景与物体的泛化,以及更快、更自然的执行,仍需要持续验证与改进。这套系统的价值,在于让每一个新问题都能更快进入数据、训练和评测。Data production, joint world–action modeling, and feedback infrastructure now form a working chain. More open-ended language following, transfer across scenes and objects, and faster, more natural execution still require continued validation and improvement. The value of this system is that every new problem can return more quickly to data, training, and evaluation.
更通用的语言跟随More general language following
开放语言指令仍要求更可靠的理解—执行一致性。Open-ended instructions still require more reliable alignment between understanding and execution.
跨场景、跨物体泛化Generalization across scenes and objects
跨陌生场景与物体的稳定迁移,仍需独立验证。Stable transfer to unfamiliar scenes and objects still needs independent validation.
更快、更自然地执行Faster, more natural execution
进一步提高执行速度与动作节奏,仍是独立的能力边界。Higher execution speed and more natural motion timing remain distinct capability boundaries.
极目于穹,洞彻至微。
我们已形成全栈的基础框架,并持续加速数据、模型与评测迭代;
期待与大家分享更多我们在通用智能上的进展。Look to the vastness, understand the subtle. With a full-stack
foundation in place, we continue to accelerate iteration across
data, models, and evaluation, and look forward to sharing more
of our progress toward general intelligence.
BibTeX
引用本文Cite this work
复制以下 BibTeX 条目即可引用本文。Use the BibTeX entry below to cite this research preview.
@misc{noematrix2026noe0,
author = {{Noematrix Team}},
title = {Noe-0 Research Preview: Breaking Through the Dexterity Bottleneck with End-to-End Non-Embodied Data},
year = {2026},
month = aug,
howpublished = {Research Blog},
url = {https://lab.noematrix.ai/blog/1-noe-0-research-preview/}
}
Noe-0 / Research Preview