< home

Embodied AI Models I’m Following

1/6 A quick map of the embodied AI model teams I’m recently following.

I’m mainly interested on the "robo brain", including model, data, and training.

Three broad groups: VLA/action policies, world-action models, and alternative architectures.

A working map of embodied AI model teams across VLA action policies, world-action modeling, and alternative architectures


2/6 VLA / action-policy models:

  • Physical Intelligence — π0.7
  • Figure — Helix 02
  • AgiBot — GO-1 / ViLLA
  • NVIDIA — GR00T
  • Google DeepMind — Gemini Robotics

Common pipeline: VLM initialization → physical pretraining → robot/task post-training → deployment feedback.


3/6 Physical pretraining may combine human video, wearable data, teleoperation, heterogeneous robot trajectories, simulation, and autonomous experience.

Disclosed scale:

  • AgiBot World: 1M+ robot trajectories
  • Figure S0: 1,000+ hours of human motion
  • PI, Google, NVIDIA: complete totals undisclosed

4/6 Post-training adapts a foundation model to a specific robot, task, and environment using demonstrations, corrections, RL, and autonomous rollouts.

Public data is limited:

  • GEN-1: ~1 hour per reported task
  • PI: some RL results use a few hours
  • Most others: undisclosed

Embodied AI training pipeline from physical data and VLM initialization through pretraining, post-training, deployment, and feedback


5/6 World and action modeling:

  • NVIDIA — Cosmos 3

VLA: predicts which action to take.

World model: predicts what may happen after an action.

WAM: jointly models future world states and actions.

Cosmos 3 spans reasoning, world generation, and action generation, so these categories overlap.

Comparison of VLA action policies, world models, and world-action models using a robotic arm task


6/6 An alternative architecture:

  • Generalist AI — GEN-1

Generalist says GEN-1 was trained from scratch on 500K+ hours of wearable human physical-interaction data, with no robot data in pretraining.

For the reported tasks, adaptation used ~1 hour of robot data per task.

Independent validation remains limited.

GEN-1 training route from wearable human physical-interaction pretraining data to robot task adaptation