2026年9月10日

ByteDance Unveils GR-3: A General-Purpose Robot “Brain” That Learns New Tasks with Minimal Data and Masters Flexible Objects

On July 22, ByteDance’s Seed team introduced GR-3, a next-generation Vision-Language-Action (VLA) model. Unlike previous systems that require massive robot trajectory datasets, GR-3 can be efficiently fine-tuned on new tasks or unseen objects with just a handful of human demonstrations. It understands abstract language commands and performs dexterous manipulation of soft objects.

1. Sample-Efficient VLA Architecture

  • Minimal Data, Maximum Transfer
    Traditional VLA models demand extensive robotic trajectory data. GR-3, by contrast, needs only 10–20 human demonstration trajectories (collected via VR) plus modest on-robot data to adapt to new environments or novel items.
  • Diverse Training Sources
    The team fused high-quality teleoperation recordings, VR-captured human motions, and large-scale publicly available vision-language datasets in joint training—enabling GR-3 to generalize far beyond earlier VLA heads like π0.

2. Unprecedented Dexterity in Real-World Tasks

  1. Robust Long-Horizon Planning
    In a multi-step “dinner table cleanup” test with 10+ subtasks, GR-3 achieved a 92% success rate, faithfully following each human-issued instruction in sequence.
  2. Soft-Object Mastery
    During a garment-hanging task, GR-3’s dual arms coordinated to pick, drape, and adjust varied clothing items with under 5% failure, even when fabrics deformed unpredictably.
  3. Understanding Abstract Commands
    Faced with “Place the unseen celadon bowl into the basket,” GR-3 executed correctly over 85% of the time—far surpassing baseline models that lack abstract reasoning.

3. ByteMini: The Agile Robotic “Body”

  • 22 Degrees of Freedom
    ByteMini integrates dual-arm manipulators, spherical wrist joints, and an omnidirectional base, allowing it to navigate tight spaces and perform Level-4 to Level-5 dexterous tasks like assembly and fine placement.
  • Rapid Deployment
    Modular hardware and native ROS compatibility enable quick integration of the GR-3 “brain,” accelerating development from prototype to production.

4. Quantified Performance Gains

  • Generalization to New Objects
    Augmenting training with public image-text data boosted GR-3’s success on novel items by 33.4%. Just 10 VR-collected trajectories lifted its baseline <60% success to >80%.
  • Instruction Robustness
    When given impossible or conflicting commands (e.g., “Put the blue bowl in the basket” when no blue bowl exists), GR-3 detects the mismatch and safely does nothing, avoiding errors.
  • Consistent Multi-Step Execution
    In a 15-step pick-and-sort scenario, GR-3 maintained under 8% failure across all stages, showcasing reliability for long sequences.

5. Looking Ahead: Toward a Universal Robot “Mind”

The Seed team envisions GR-3 as a stepping stone to truly general-purpose robotic intelligence: pre-trained on diverse multimodal data, and fine-tuned rapidly with few examples. Future deployments may span industrial assembly, warehouse sorting, medical assistance, and home service—fulfilling the promise of a versatile, abstract-reasoning robotic “brain.”

接著讀

ChatGPT: Everything you need to know about the AI-powered chatbot

ChatGPT, OpenAI’s text-generating AI chatbot, has taken the world by storm since its launch in November 2022. What started as a tool to supercharge productivity through writing essays and code with short text prompts has evolved into a behemoth with 300 million weekly active users.

496 天前