NOEMATRIX LAB研究博客 NO.1Research Blog NO.1

Noe-0研究概览 Noe-0 Research Preview

全流程无本体数据,
突破灵巧性瓶颈。
End-to-end non-embodied data.
Breaking through the dexterity bottleneck.

目前机器人的后训练通常依赖遥操作,轨迹中的停顿、分段与反复修正可能压缩预训练阶段获得的动作知识。Noe-0 在预训练与后训练中统一使用无本体数据,使自然动作直接迁移为机器人任务能力。 Conventional post-training often relies on robot teleoperation. Pauses, segmentation, and repeated corrections in teleoperated trajectories can narrow the motion knowledge acquired during pre-training. Noe-0 uses non-embodied data in both stages, allowing natural motion knowledge to transfer more directly into robotic task capabilities.

Noe-0 自主执行多类真实任务的完整演示集锦;视频以原速播放。 A longer reel of Noe-0 executing diverse real-world tasks autonomously at the original playback speed.

引言Introduction

从零构建数据、模型与基础设施的全栈闭环。A full-stack loop built from scratch across data, models, and infrastructure.

01引言 / 全栈系统Introduction / Full-stack system

具身基础模型,依赖全栈系统。Embodied foundation models demand a full-stack system.

具身通用基础模型不是单一的模型架构问题,而是一项全栈系统工程。真实世界的数据生产、物理世界与动作的联合建模,以及支撑规模化迭代的基础设施,都需要从零开始、协同建设,才能让数据、训练与反馈形成高效、完整的闭环。Building a general-purpose foundation model for embodied intelligence is not solely a model-architecture problem; it is a full-stack systems problem. Real-world data production, joint world–action modeling, and infrastructure for iteration at scale must all be built from scratch and developed together, so that data, training, and feedback form an efficient, complete loop.

经过过去数月的工作,我们很高兴在这篇研究概览中分享 Noe-0 的近期进展与当前验证结果。这里展示了已经完成的关键系统环节:无本体数据的采集与治理、世界—动作联合建模,以及连接数据生产与管理的基础设施。我们也附上了一些真实数据样例、实际的系统流程和训练曲线,共同呈现这条路径目前已经走到哪里。After several months of work, we are pleased to share Noe-0’s recent progress and the results validated to date. This Research Preview presents the key parts of the system now in place: non-embodied data collection and governance, joint world–action modeling, and infrastructure connecting data production with natural-language querying. Together, the real-world samples, system workflows, and training curves show how far this path has progressed.

01

数据

Data

真实世界采集、整理与版本化。

Capture, curation, and versioning in the real world.

02

模型

Model

联合学习世界变化与动作。

Jointly learn world dynamics and action.

03

基础设施

Infrastructure

让每次处理、训练与验证可追踪。

Keep processing, training, and validation traceable.

Figure 01 Noe-0 将真实数据生产、世界—动作联合建模与可追踪迭代连接成同一反馈系统。Noe-0 couples real-world data production, joint world–action modeling, and traceable iteration into one feedback system.

数据Data

从自然动作到可扩展数据。From natural motion to scalable data.

02.A数据 / 自然动作Data / Natural motion

动作质量,始于采集。Motion quality begins with how it is captured.

人类在操作过程中,手部会连续调整姿态、接触力与节奏。RoboPocket 使用同步的头部与双腕视角,在不依赖机器人遥操作的前提下,记录这种自然、灵巧的操作过程。Human hands continuously adjust pose, contact, and rhythm. RoboPocket uses synchronized head-mounted and dual-wrist views to capture natural dexterous behavior without robot teleoperation.

RoboPocket / natural capture

保留动作,再形成能力。Preserve motion. Build capability.

三个同步视角同时记录手、物体与环境变化:头部视角保留任务意图,两路手腕视角补充接触与局部动作。Three synchronized views capture hands, objects, and scene change together: the head view preserves task intent while two wrist views reveal contact and local motion.

预训练从这些无本体数据中学习细粒度手部运动、物体状态变化与动作节奏;它看到的不是遥操作能够复现的机械性的运动轨迹,而是人类操作本身更自然的动作分布。During pre-training, the model learns fine-grained hand motion, object-state changes, and temporal dynamics from this non-embodied data. This exposes the model to a broader distribution of human manipulation than the mechanically constrained trajectories typically produced by teleoperation.

如果后训练重新换成停顿、机械性动作和频繁修正的遥操作轨迹,早期学到的动作知识就会被压缩到更窄的分布中。Noe-0 在两个阶段沿用无本体数据,让自然动作的灵巧性可以完整保留在模型中。If post-training switches back to hesitant, mechanically constrained, and frequently corrected teleoperation trajectories, earlier motion knowledge is compressed into a narrower distribution. Noe-0 uses non-embodied data during both stages to reduce the loss of natural motion characteristics during post-training.

后训练数据分布影响预训练动作知识的保留与迁移 Post-training data shapes how pre-training motion knowledge is retained and transferred 对比遥操作后训练与无本体数据后训练中的动作分布连续性。 Motion-distribution continuity under teleoperated and non-embodied post-training. 常规路径 · 遥操作后训练 CONVENTIONAL · TELEOPERATED POST-TRAINING 自然动作分布 natural motion distribution 预训练 Pre-training 无本体数据 non-embodied data 遥操作后训练 TELEOPERATED POST-TRAINING 灵巧性瓶颈 DEXTERITY BOTTLENECK 分段 · 停顿 · 反复修正 segmented · hesitant · corrective NOE-0 · 全流程无本体数据 NOE-0 · END-TO-END NON-EMBODIED DATA 自然动作分布 natural motion distribution 预训练 Pre-training 无本体数据 non-embodied data 后训练 Post-training 分布连续 CONTINUOUS 无本体数据 non-embodied data 连续 · 自然 · 灵巧 continuous · natural · dexterous 常规路径 · 遥操作后训练 CONVENTIONAL · TELEOPERATED POST-TRAINING 动作分布MOTION 预训练PRETRAIN无本体数据NON-EMBODIED 灵巧性瓶颈DEXTERITY BOTTLENECK 分段 · 停顿 · 修正SEGMENTED · HESITANT · CORRECTIVE NOE-0 · 全流程无本体数据 NOE-0 · END-TO-END NON-EMBODIED DATA 动作分布MOTION 预训练PRETRAIN无本体数据NON-EMBODIED 后训练POST-TRAIN分布连续CONTINUOUS 连续 · 自然 · 灵巧CONTINUOUS · NATURAL · DEXTEROUS 常规路径 · 遥操作后训练CONVENTIONAL · TELEOPERATED 预训练PRETRAIN无本体NON-EMBODIED 灵巧性瓶颈DEXTERITYBOTTLENECK 自然动作分布被压缩为停顿、分段与反复修正的机械性动作。Natural motion is compressed into segmented,repeatedly corrected mechanical motion. NOE-0 · 全流程无本体数据NOE-0 · END-TO-END NON-EMBODIED 预训练PRETRAIN无本体NON-EMBODIED 后训练POST连续CONTINUOUS 分布连续CONTINUOUSDISTRIBUTION 同一种自然动作贯穿预训练与后训练。One natural motion pattern remains continuousacross stages.
Figure 02常规遥操作后训练可能将自然动作知识压缩为更机械性的运动轨迹;Noe-0 在预训练与后训练中统一使用无本体数据,以保持动作分布的连续性与自然性。Conventional teleoperated post-training can compress natural motion knowledge into more mechanical trajectories. Noe-0 uses non-embodied data in both pre-training and post-training to maintain a continuous, natural motion distribution.

采集自然动作Capture natural motion

头部与双腕同步记录任务意图、局部接触和连续动作。Head and dual-wrist views preserve intent, contact, and continuous motion.

突破遥操作瓶颈Break through the teleoperation bottleneck

后训练不再用机械化轨迹重新限制预训练知识。Post-training does not re-limit pre-training knowledge with mechanical trajectories.

保持动作语言连续Keep one action language

同一种无本体动作分布从预训练延续到任务能力形成。One non-embodied action distribution continues into task-level capability.

02.B数据 / 真实场景规模Data / Real-world scale

集中式数据采集受限于标准化搭建的采集场景,操作对象大多批量统一采购,因此难以自然获得足够的数据多样性。若将数据采集下沉至采集人员的生活空间,环境内日积月累的包装、工具、容器和日用品,可覆盖材质、形态及使用状态下的各类长尾样本。Centralized data collection is constrained by standardized collection environments, where objects are typically purchased in uniform batches, making sufficient diversity difficult to obtain naturally. Moving collection into collectors’ living spaces brings in accumulated packaging, tools, containers, and everyday items, covering long-tail samples across materials, forms, and conditions of use.

47城市cities
18省级行政区provincial regions
2000+平台累计支持采集人员cumulative collectors supported
累计数据时长Cumulative data hours 单位 / 小时Unit / hours
Cumulative data hours Four verified cumulative totals: 4,761.2, 19,650.1, 61,151.9, and 97,387.7 hours. No calendar dates are shown. 4,761.2 19,650.1 61,151.9 97,387.7 h CURRENT
Figure 03累计数据时长,最新汇总为 97,387.7 小时。Cumulative data hours, with the latest total at 97,387.7 hours.
部分数采人员城市分布 Selected collector city distribution 47 城 · 18 个省级行政区。 47 cities · 18 provincial-level regions. 代表性任务类型TASK TYPES 新乡Xinxiang 郑州Zhengzhou 青岛Qingdao 徐州Xuzhou 盐城Yancheng 扬州Yangzhou 南京Nanjing 滁州Chuzhou 合肥Hefei 马鞍山Maanshan 芜湖Wuhu 杭州Hangzhou 池州Chizhou 武汉Wuhan 吉安Ji'an 衢州Quzhou 赣州Ganzhou 01 / 工具测量01 / TOOL MEASUREMENT 02 / 纸张整理02 / PAPER HANDLING 03 / 电子部件装配03 / DEVICE ASSEMBLY
01郑州Zhengzhou
02合肥Hefei
03武汉Wuhan
Head and dual-wrist views of measuring a power bank with a voltage meter

工具测量

Tool measurement

Head and dual-wrist views of rolling a Spring Festival couplet

纸张整理

Paper handling

Head and dual-wrist views of assembling a game controller

电子部件装配

Device assembly

Figure 04部分代表性的采集覆盖城市,完整采集网络覆盖 47 个城市、18 个省级行政区。省界分界线为示意图。Representative cities covered by data collection are shown; the full network spans 47 cities across 18 provincial-level regions. Provincial boundaries are schematic.

02.C / 真实世界数据02.C / REAL-WORLD DATA

真实世界数据样例。Real-world data samples.

覆盖精细装配、工具使用、柔性物体、分类归置与多阶段操作。点击按钮可随机查看另一组数据样例。Spanning precision assembly, tool use, deformable objects, sorting, and multi-stage manipulation. Use the button to view another selection.

模型Model

同时进行隐式世界想象和工作规划。Perform implicit world imagination and task planning simultaneously.

03.AWorld × Action

一边隐式想象世界,
一边规划下一步动作。
Imagine the world implicitly while planning the next action.

Noe-0 接收图像与任务语义后,在同一个表示中同时进行隐式世界想象与动作规划,并输出动作。它并非“先想象、再规划”的串行流水线,而是在行动发生时持续保持对世界变化的预测。Given images and task semantics, Noe-0 performs implicit world imagination and action planning in one shared representation while producing actions. It is not a serial imagine-then-plan pipeline; prediction of world change remains active as actions unfold.

隐式想象为行动规划提供对象状态与因果上下文;行动信号又反过来约束模型,只关注真正影响任务进程的变化。两者共享时间步与表示空间,而非通过前后两个独立模块顺序传递。Implicit imagination supplies object state and causal context to action planning; action signals in turn constrain the model to changes that matter to task progress. The two share a time step and representation space rather than passing through separate modules in sequence.

执行动作后,新观察进入下一轮联合推理。完整闭环是“观察 → 隐式想象与行动规划同时发生 → 执行动作 → 再观察”,从而在长时序任务中持续更新世界状态与控制决策。After the action executes, a new observation enters the next joint inference step. The full loop is observe → imagine implicitly and plan simultaneously → execute → observe again, continuously updating world state and control decisions across long-horizon tasks.

Noe-0 joint world-action inference Images and task semantics enter a shared representation. Noisy action tokens can attend to noisy video tokens while world imagination and action planning run in parallel. 04 / WORLD × ACTION 图像与任务语义进入共享表示,隐式世界想象与动作规划并行发生。 Images and task semantics enter a shared representation for parallel world and action reasoning. 01 / INPUT 图像 Images 任务语义 Task semantics 真实头部相机画面 real head-camera image Noe-0 世界—动作联合去噪 joint world–action denoising 噪声动作 → 噪声视频注意力 NOISY ACTION → NOISY VIDEO 隐式世界想象 Implicit world imagination NOISY VIDEO TOKENS 预测与任务相关的世界变化 predict task-relevant change 动作规划 Action planning 同一共享表示 same shared representation aₜ +1 +2 +H NOISY ACTION TOKENS 在预测世界变化时同步输出 emit actions while predicting 共享表示 · 同一时间步 · 相互约束 shared representation · same timestep · mutual constraint 03 / ACTION + FEEDBACK 动作输出 + 观察反馈 Action + observation 动作轨迹 ACTION TRAJECTORY
01

图像 + 任务语义

Images + task semantics

02

Noe-0

噪声动作 Token 可见噪声视频 Token;隐式世界想象与动作规划并行发生。

Noisy action tokens attend to noisy video tokens while world imagination and action planning run in parallel.

03

动作输出 + 下一次观察

Action output + next observation

Figure 05真实图像与任务语义进入 Noe-0。噪声动作 Token 可对噪声视频 Token 进行注意力计算,使隐式世界想象与动作规划在共享表示的同一时间步内相互约束;动作执行后的新观察进入下一轮联合推理。Real images and task semantics enter Noe-0. Noisy action tokens attend to noisy video tokens, coupling implicit world imagination and action planning in one shared representation and timestep. The observation after execution enters the next joint-inference step.
03.B模型 / 训练证据Model / Training evidence

用模型扩展,
验证数据有效性。
Use scaling behavior in model training to validate data effectiveness.

在固定的验证集设置下,10%、20% 与 100% 数据量实验呈现 loss 的有序下降。Under a fixed validation-set protocol, the 10%, 20%, and 100% data-volume runs show an ordered decline in loss.

验证动作损失 / 越低越好Validation action loss / lower is better固定验证集Fixed validation set
.045 .042 .039 .036 56k 70k 84k 100k 优化器步数 / 每 2k 评测 optimizer steps / evaluated every 2k 100% data 20% data 10% data 10% · 0.0439 20% · 0.0403 100% · 0.0364

固定验证集 / 越低越好FIXED VALIDATION SET / LOWER IS BETTER

10% DATA0.043972k
20% DATA0.040382k
100% DATA0.0364100k
Figure 06在固定验证集设置下,曲线显示每 2k 步记录的评测点;1.4 万小时子数据集上的三个实验仅改变数据量。Under one fixed validation set, the curves show evaluation points recorded every 2k steps. Only data volume changes across the three runs, all based on a 14,000-hour subdataset.
01 / TRACE

完整评测轨迹Complete traces

从 56k 步开始,每 2k 步记录一次验证动作损失。Validation action loss is recorded every 2k steps from 56k onward, so endpoints never stand in for the process.

02 / VALIDATION

固定验证集Fixed validation set

三组实验使用同一固定验证集设置。All three runs use the same fixed validation set.

03 / VARIABLE

仅改变数据量Data volume only

三组实验仅改变训练数据量。Only training-data volume changes.

基础设施Infrastructure

让数据生产与模型迭代形成可追踪闭环。A traceable loop between data production and model iteration.

04.A基础设施 / 数据生产Infrastructure / Data production

从端侧采集,
到可交付的训练资产。
From edge capture to training-ready assets.

端侧数采系统 RoboPocket 的生产链把端侧三视角采集、任务管理、采集工具、数据管理平台 DM3 云端入库、质量检查、人工审核、标注审核与数据交付连成一条可追踪链路。The edge data-collection system RoboPocket connects three-view edge capture, task management, capture tooling, cloud ingest into the data management platform DM3, quality checks, human and annotation review, and data delivery in one traceable path.

端侧负责把同步、完整的多视角任务数据带回云端;云端连续完成入库、自动检查与人工判断。每一次质量结论都会与任务、采集片段和数据版本保持关联,便于回溯与修正。The edge returns synchronized, complete multi-view task data; the cloud performs ingest, automated checks, and human review. Every quality decision stays linked to its task, capture, and data version for traceability and revision.

在已审批数据切片的同口径对照中,半自动标注用时较纯人工减少约 22%;自动化处理高频判断,人处理难例;交付后的训练资产继续与配置、检查点和评测反馈关联,把模型中暴露的问题送回下一轮数据工作。In a like-for-like comparison on an approved data slice, semi-automated annotation reduced time by approximately 22%; Automation handles recurring decisions while people resolve hard cases. Delivered assets remain linked to training configurations, checkpoints, and evaluation feedback, returning model failures to the next data cycle.

从端侧采集到数据交付,每个数据版本与质量结论均可追踪 From edge capture to delivery, every data version and quality decision remains traceable Data Agent 通过自然语言协同数采管理员与数采员。 The Data Agent connects data collection managers and data collectors through natural-language interaction. 01 端侧采集Edge capture 头部 + 双腕同步head + dual-wrist sync MULTI-VIEW / SYNC 02 任务管理Task management 任务定义 · 边界task spec · scope TASK SPEC 03 采集工具Capture tool RoboPocket CAPTURE / UPLOAD 04 云端入库Cloud ingest DM3 UPLOAD / INDEX 05 云端质检Cloud QC 同步 · 完整性sync · integrity AUTOMATED QC 06 人工审核Human review 困难案例 · 修正hard cases · correction REVIEW / TRACE 07 标注审核Annotation review AI 预标注 + 人工优化AI pre-labeling + human refinement LABEL / VERIFY 08 数据交付Data delivery 语义聚类 + 数据筛选clustering + filtering DATASET DELIVERY DATA AGENT / NATURAL-LANGUAGE OPERATIONS 数采管理员Data collection managers TASK / POLICY Data Agent 任务生成 · 查询统计 · 质量问题处理 · 消息通知 · 权限管理 tasking · query · issue resolution · notifications · access control TRACEABLE COORDINATION 数采员Data collectors CAPTURE / FEEDBACK
01

端侧采集

Edge capture

头部 + 双腕同步

head + dual-wrist sync

02

任务管理

Task management

03

采集工具

Capture tool

RoboPocket

04

云端入库

Cloud ingest

DM3

05

云端质检

Cloud QC

06

人工审核

Human review

07

标注审核

Annotation review

AI 预标注 + 人工优化

AI pre-labeling + human refinement

08

数据交付

Data delivery

语义聚类 + 数据筛选

clustering + filtering

Data Agent

以自然语言连接数采管理员与数采员。

A natural-language bridge between data collection managers and data collectors.

Figure 07端侧数采系统 RoboPocket 与数据管理平台 DM3 将三视角数据送入审核、筛选和交付;Data Agent 连接数采管理员与数采员。The edge data-collection system RoboPocket and data management platform DM3 carry three-view data through review and delivery; the Data Agent connects managers and collectors.

04.B / Data Agent

把分散的数据操作,
组织成一条可追踪工作流。
Bring distributed data operations into one traceable workflow.

Data Agent 连接任务生成、数据查询与统计、质量问题处理、消息通知和权限管理。用户可以直接用自然语言查询采集进度、任务状态与质量问题。它不替代人的质量判断,而是把分散的操作组织成同一条可回溯工作流,让每次修正直接服务于后续数据版本。The Data Agent connects task generation, query and statistics, quality issue resolution, notifications, and access control. Users can query collection progress, task status, and quality issues directly in natural language. It does not replace human quality judgment; it organizes scattered operations into a traceable workflow so every correction reaches later data versions.

任务生成Task generation 查询与统计Query & stats 质量问题处理Quality issue resolution 消息通知Notifications 权限管理Access control
完整的 Data Agent 中文横版聊天界面,展示任务质检记录、自然语言查询、采集概览与质量反馈
Figure 08 Data Agent 通过自然语言查询汇总采集进度,并将质量反馈传递给一线采集员。The Data Agent turns natural-language queries into collection summaries and routes quality feedback to frontline collectors.
Data layer

让重复判断交给半自动链路Make recurring judgment semi-automatic

任务理解、质量检查、预标注与置信度判断形成统一入口,让重复判断进入稳定的半自动链路。Task interpretation, quality review, pre-annotation, and confidence estimation share one entry point, moving recurring decisions into a stable semi-automated workflow.

Training layer

让训练资产可追踪、可复用Keep training assets traceable and reusable

数据版本、训练配置、检查点与评测结果保持关联,使每次迭代都有完整依据。Data versions, training configurations, checkpoints, and evaluations remain linked so every iteration has a complete evidence trail.

Learning loop

让每轮训练改善下一轮数据Let every model improve the next dataset

模型反馈驱动任务补齐与采集修正,使数据生产与能力增长形成闭环。Model feedback drives task coverage and capture correction, closing the loop between data production and capability growth.

展望Outlook

闭环不是终点,而是更快逼近通用性的起点。The loop is not the finish line; it is the engine for approaching generality faster.

05展望 / 通向通用智能Outlook / Toward general intelligence

闭环迭代,持续突破能力边界。Close-loop iteration, continuously pushing the boundaries of capability.

目前,数据生产、世界—动作联合建模与反馈基础设施已经连接成一条可工作的链路。至于更开放的语言跟随、跨场景与物体的泛化,以及更快、更自然的执行,仍需要持续验证与改进。这套系统的价值,在于让每一个新问题都能更快进入数据、训练和评测。Data production, joint world–action modeling, and feedback infrastructure now form a working chain. More open-ended language following, transfer across scenes and objects, and faster, more natural execution still require continued validation and improvement. The value of this system is that every new problem can return more quickly to data, training, and evaluation.

01 / LANGUAGE

更通用的语言跟随More general language following

开放语言指令仍要求更可靠的理解—执行一致性。Open-ended instructions still require more reliable alignment between understanding and execution.

02 / GENERALIZATION

跨场景、跨物体泛化Generalization across scenes and objects

跨陌生场景与物体的稳定迁移,仍需独立验证。Stable transfer to unfamiliar scenes and objects still needs independent validation.

03 / SPEED

更快、更自然地执行Faster, more natural execution

进一步提高执行速度与动作节奏,仍是独立的能力边界。Higher execution speed and more natural motion timing remain distinct capability boundaries.

极目于穹,洞彻至微。
我们已形成全栈的基础框架,并持续加速数据、模型与评测迭代;
期待与大家分享更多我们在通用智能上的进展。
Look to the vastness, understand the subtle. With a full-stack foundation in place, we continue to accelerate iteration across data, models, and evaluation, and look forward to sharing more of our progress toward general intelligence.

06引用Citation

BibTeX

引用本文Cite this work

复制以下 BibTeX 条目即可引用本文。Use the BibTeX entry below to cite this research preview.

@misc{noematrix2026noe0,
  author       = {{Noematrix Team}},
  title        = {Noe-0 Research Preview: Breaking Through the Dexterity Bottleneck with End-to-End Non-Embodied Data},
  year         = {2026},
  month        = aug,
  howpublished = {Research Blog},
  url          = {https://lab.noematrix.ai/blog/1-noe-0-research-preview/}
}

Noe-0 / Research Preview

全流程无本体数据,
突破灵巧性瓶颈。
End-to-end non-embodied data.
Breaking through the dexterity bottleneck.

返回文章开头Back to the beginning