Skip to content
@AMAP-ML

DreamX

DreamX — AMAP Spatial Intelligence

DreamX

AMAP Spatial Intelligence Models & Systems

Understand · Predict · Simulate · Create · Plan · Act in the real world.

关注 AMAP-ML 探索 projects Join DreamX

DreamX is the spatial intelligence portfolio of Alibaba AMAP, built and released by the DreamX team through AMAP-ML.
We connect research, production systems, and real-world deployment across maps, mobility, local services, digital content, and interactive worlds.
The team releases open-source projects, benchmarks, and publications at ICLR, CVPR, ECCV, ACL, AAAI, SIGGRAPH, SIGGRAPH Asia, IJCV, ICCV, ICML, KDD, CIKM, EMNLP, ACM MM, and WWW.

Flagship releases · Research directions · Recent updates · All projects · Collaborate with us


At a glance

40+ open-source projects  ·  52 papers at major venues in 2026  ·  3 core problems  ·  6 DreamX model and system families

Understand & Predict Generate & Simulate Plan & Act
Maps, mobility, places, intent, and urban dynamics World models, media, scenes, and digital assets Agents, tool use, decision-making, and embodied systems
DreamX-Predictor · DreamX-REC DreamX-World · DreamX-Creator DreamX-Agent · DreamX-Phi

Flagship releases

Project What it brings to spatial intelligence Links
DreamX-World 1.0 Interactive, long-horizon world simulation with an open 5B model. Code · Report
DreamX-Phi 1.0 Geometry-aware, action-conditioned video world modeling for robotic manipulation. Code · Paper
LongHorizon-Harness Durable task state and independent verification for reliable computer-use agents. Code · Paper
MobilityBench Real-world benchmark for route-planning agents in AMAP-native mobility scenarios. Code
SkillClaw Evolves reusable agent skills from real interaction traces. Code
FluxText Controllable scene-text editing for practical visual-asset creation. Code

Browse the complete project map →


Research directions

A shared mission: spatial intelligence

We build AI that understands the real world and its change over space and time, predicts what comes next, generates and simulates digital counterparts, plans toward human goals, and acts through products and embodied systems.

For AMAP, this closes a learning-and-deployment loop across maps, routes, places, mobility, urban environments, local services, digital content, and physical action.

Direction DreamX families Representative work
Understand and predict the world DreamX-Predictor · DreamX-REC MobilityBench, Thinking-with-Map, SocioReasoner
Generate and simulate the world DreamX-World · DreamX-Creator DreamX-World, OmniDance, RL3DEdit
Plan and act in the world DreamX-Agent · DreamX-Phi DreamX-Phi, AutoDrive-R2, GPG

One foundation, six families

All DreamX families share spatial data and knowledge, multimodal foundation models, spatiotemporal and world modeling, agent and reinforcement-learning methods, generative modeling, plus scalable infrastructure and evaluation.

Family Role
DreamX-Predictor Models the evolution of traffic, mobility, demand, supply, and urban conditions.
DreamX-World Learns controllable, persistent, physically grounded, interactive world models.
DreamX-Agent Reasons over spatial context, uses tools, and completes map, mobility, and local-service workflows.
DreamX-Phi Connects perception, reasoning, decision-making, and physical action.
DreamX-REC Matches intent with places, content, routes, and services under real-world constraints.
DreamX-Creator Generates and edits map assets, local-service content, images, videos, 3D scenes, and spatial media.

Recent updates

  • 2026.08.21 — The DreamX team has 3 papers accepted to EMNLP 2026, advancing multimodal continual learning, self-evolving agents, and language-model reasoning.
  • 2026.08.13DreamX-Phi 1.0 launches a geometry-aware, action-conditioned video world model for robotic manipulation, leading WorldArena 2.0 Track 1 in the August 12 leaderboard snapshot. Paper
  • 2026.08.03LongHorizon-Harness introduces a Manage–Execute–Audit loop for durable, verified computer-use progress. Paper
  • 2026.07.20OmniDance is selected as an ECCV 2026 Oral, advancing large-scale multimodal dance-video generation from text, image, and music.
  • 2026.06.18 — Five DreamX papers are accepted to ECCV 2026, spanning spatial intelligence, generative modeling, and multimodal AI.
  • 2026.06.15DreamX-World releases its 1.0 technical report and the DreamX-World-5B model for one-minute interactive world generation. Report
Earlier updates

Project map

Understand and predict the world
仓库 Contribution Venue
MobilityBench Route-planning agent evaluation in real-world mobility scenarios. KDD 2026 Oral
Thinking-with-Map Map-augmented geolocalization agent trained with reinforcement learning. ACL 2026 Findings
SocioReasoner Vision-language reasoning for urban socio-semantic segmentation. ICLR 2026
DSFNet Multi-scenario route ranking with a public industrial driving-route dataset and AMAP deployment. WWW 2025
GenMRP Generative multi-route planning for efficient, personalized, real-time industrial navigation. CIKM 2026
IntSR Integrated generative modeling for search and recommendation across AMAP scenarios. CIKM 2026
IntTravel Real-world dataset and generative framework for integrated multi-task travel recommendation. arXiv 2026
FE2E Image-editing priors for dense geometry estimation. CVPR 2026
UniVG-R1 Reasoning-guided universal visual grounding with reinforcement learning. arXiv 2025
Taming-Hallucinations Counterfactual video generation for reducing MLLM video hallucinations.
Generate and simulate the world
仓库 Contribution Venue
DreamX-World 1.0 General-purpose world model for interactive world simulation.
Code2World GUI world model via renderable code generation.
FluxText Diffusion transformer baseline for scene-text editing.
OmniDance Multimodal dance-video generation from text, image, and music. ECCV 2026 Oral
SCALAR / SCALAR++ Efficient controllable generation through scale-wise visual autoregressive learning. AAAI 2026 / IJCV
MAR-GRPO Stabilized RL for autoregressive-diffusion image generation. ACM MM 2026
RL3DEdit Geometry-guided RL for multi-view-consistent 3D scene editing. ECCV 2026
MACE-Dance Motion–appearance cascaded generation for music-driven dance video. SIGGRAPH 2026
Omni-Effects Prompt-guided, spatially controllable composite visual-effects generation. AAAI 2026
S2-Guidance Training-free stochastic self-guidance for diffusion models. ICLR 2026
EPG Pixel-space generative modeling via self-supervised pre-training. ICLR 2026
USP Unified self-supervised pretraining in VAE space for diffusion models. ICCV 2025
EMF Text-conditioned one-step image generation. CVPR 2026
DCW Differential correction for SNR-t bias in diffusion probabilistic models. CVPR 2026
NarrLV Narrative-centric evaluation for long-video generation models. ICLR 2026
Imagery搜索 Adaptive test-time search for video generation. AAAI 2026
Eevee High-resolution benchmark for video-based virtual try-on. CVPR 2026 Findings
VMBench Perception-aligned benchmark for video motion generation. ICCV 2025
Plan and act in the world
仓库 Contribution Venue
DreamX-Phi 1.0 Geometry-aware, action-conditioned video world modeling for bimanual robotic manipulation. arXiv 2026
LongHorizon-Harness Verified long-horizon computer use through durable task state and Manage–Execute–Audit loops. arXiv 2026
SkillClaw Agentic evolver for collective skill-library improvement.
AutoDrive-R2 Reasoning and self-reflection for VLA models in autonomous driving. ICLR 2026
Tree-GRPO Tree-search rollouts for LLM-agent reinforcement learning. ICLR 2026
GPG Simple and strong group policy-gradient baseline for model reasoning. ICLR 2026
CoEvolve Agent–data mutual evolution for training LLM agents. ACL 2026
Role-Agent Dual-role evolution that trains an LLM as both an agent and an environment model. EMNLP 2026
MathForge Difficulty-aware GRPO and multi-aspect reformulation for math reasoning. ICLR 2026
Video-STAR Tool-using RL for open-vocabulary action recognition. ICLR 2026
Shared foundations and evaluation
仓库 Contribution Venue
SpatialGenEval Spatial-intelligence evaluation for text-to-image models. ICLR 2026
Omni-WorldBench Benchmark for interactive response capabilities of world models. arXiv 2026
RealQA Realistic image-quality and aesthetic scoring with multimodal LLMs.
M2Note Training-free continual evolution for vision-language models through an editable mistake notebook. EMNLP 2026
Peak-End-Net Peak-end-rule-inspired framework for generalizable video aesthetic assessment. ACM MM 2026
Evaluation-Verification Reward Fine-grained reward modeling for consistent multi-reference image editing. SIGGRAPH Asia 2026
FASA Frequency-aware sparse attention for efficient sparse decoding. ICLR 2026

Collaborate with us

We welcome research interns, full-time researchers, AI engineers, and academic collaborators working on spatial intelligence, LLM agents, reinforcement learning, world models, multimodal learning, embodied AI, recommendation, and generative AI.

Please send your CV, representative work, and research interests to cxxgtxy@gmail.com. You can also visit our organization page or team homepage.

Pinned Loading

  1. Tree-GRPO Tree-GRPO Public

    [ICLR 2026] Tree 搜索 for LLM Agent Reinforcement Learning

    Python 399 40

  2. Code2World Code2World Public

    Code2World: A GUI World Model via Renderable Code Generation

    Python 310 16

  3. GPG GPG Public

    [ICLR26]GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning

    Python 181 5

  4. DreamX-World DreamX-World Public

    DreamX-World: A General-Purpose Interactive World Model

    Python 762 50

  5. SkillClaw SkillClaw Public

    Let Skills Evolve Collectively with Agentic Evolver

    Python 2.5k 250

  6. FluxText FluxText Public

    Implementation of "FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing"

    Python 921 32

仓库

Showing 10 of 46 repositories

Top languages

Loading…

Most used topics

Loading…