Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LLaDA-UI

Bringing Block-wise Diffusion to Vision-Language GUI Agents

Paper Project Page Hugging Face Model GitHub Code

Diffusion Model Multimodal LLM GUI Agent Mixture of Experts

LLaDA-UI is an MoE-based, block-wise diffusion vision-language GUI agent. It combines a native dynamic-resolution vision encoder with the LLaDA2.0-mini-base diffusion language backbone, and generates reasoning and GUI actions through the same block-wise diffusion decoder.

LLaDA-UI GUI-agent benchmark comparison and animated qualitative diffusion decoding

Figure 1. LLaDA-UI GUI-agent performance and qualitative block-wise diffusion decoding. The radar chart and GUI observation remain static while the model output is progressively denoised.

Highlights

  • Diffusion-native multimodal agent: LLaDA-UI extends a masked diffusion language model to visual understanding and executable GUI interaction.
  • MoE architecture: approximately 16.7B total parameters end to end.
  • Two-stage training: large-scale multimodal pre-training aligns a native-resolution SigLIP-initialized ViT with LLaDA2.0-mini-base, followed by GUI-agent supervised fine-tuning.
  • Cross-platform interaction: training covers grounding, mobile, desktop, and web tasks, including data from more than 100 Chinese mobile applications and more than 70 English mobile applications.
  • Competitive GUI performance: LLaDA-UI outperforms Qwen2.5-VL-7B across all six reported GUI benchmarks and surpasses Qwen3-VL-8B on ScreenSpot-Pro, AndroidWorld, MobileWorld, and WebVoyager.
  • Diffusion-specific analysis: the accompanying report studies structured action validity, trajectory-length effects, EOS handling, denoising steps, and block size.

Model Overview

Item Description
Model type MoE block-wise diffusion vision-language GUI agent
Total parameters Approximately 16.7B end-to-end
Language backbone LLaDA2.0-mini-base
Vision encoder Native-resolution ViT initialized from SigLIP, with 2D RoPE
Vision-language connector Spatial 4-to-1 feature grouping followed by a two-layer MLP projector
Output Text reasoning and structured GUI actions
Spatial convention Normalized coordinates in [0, 999]
Training stages Multimodal pre-training, then GUI-agent SFT
Supported domains GUI grounding, mobile, desktop, and web

Training and Inference Pipelines

Inference Pipeline

LLaDA-UI multimodal block-wise diffusion inference pipeline

Figure 2. GUI observations from web, mobile, and desktop environments are encoded at native resolution and combined with task and interaction-history tokens. The LLaDA2.0 decoder progressively denoises the model output into reasoning and executable actions.

GUI Data Generation

GUI data-generation pipeline

Figure 3. Overview of the GUI data-generation pipeline, from task construction and sub-skill decomposition to compositional trajectory collection and quality control.

Benchmark Results

GUI-Agent Evaluation

LLaDA-UI GUI-agent benchmark results from Table 3 of the technical report

GUI-agent evaluation reproduced directly from Table 3 of the technical report.

Evaluation versions, task counts, prompts, reset policies, action limits, serving backends, and baseline provenance will be documented in the final release.

Qualitative GUI-Agent Traces

The following GIFs replay complete successful trajectories. Every frame shows the original screenshot together with the verbatim model output for that step.

WebVoyager: constrained flight search

Complete WebVoyager Google Flights trajectory

LLaDA-UI configures a one-way Calgary-New York flight, navigates to December 3, 2026, applies the lower-emissions constraint, and returns the verified lowest-CO2 itinerary over a 20-step trajectory.

OSWorld: Calc to Writer

Complete OSWorld Calc-to-Writer trajectory

LLaDA-UI selects a formatted table in LibreOffice Calc, transfers it to Writer, and saves the resulting document as price.docx on the desktop.

MobileWorld: email to alarm

Complete MobileWorld email-to-alarm trajectory

LLaDA-UI reads the 7:00 PM Christmas-party time from email, computes the one-hour offset, opens the Clock application, and verifies the enabled 6:00 PM alarm.

Quickstart

Use separate environments for standalone Hugging Face inference and SGLang serving because their tested Transformers and PyTorch versions differ.

Hugging Face inference

inference/inference_hf.py is a self-contained grounding entry point; it does not require the training repository.

pip install torch==2.5.1 torchvision \
  --index-url https://download.pytorch.org/whl/cu124
pip install transformers==4.51.0 Pillow numpy einops accelerate \
  sentencepiece protobuf safetensors
pip install ninja
pip install flash-attn==2.7.4.post1 --no-build-isolation --no-cache-dir

Download the checkpoint from the LLaDA-UI Hugging Face repository, then pass its local directory to --ckpt:

CUDA_VISIBLE_DEVICES=0 IMAGE_MAX_PIXELS=12845056 \
python -u inference/inference_hf.py \
  --ckpt /path/to/hf_ckpt_unpacked \
  --image /path/to/screenshot.png \
  --prompt "close this window" \
  --gen-length 32 \
  --steps 32 \
  --block-length 32

The script prints the verbatim generation and parsed normalized point. Use IMAGE_MAX_PIXELS=1003520 for the default-resolution profile. ScreenSpot-V2 pipeline verification is also available; see python inference/inference_hf.py --help.

SGLang serving

inference/sglang_client.py is a single-file, standard-library client for an OpenAI-compatible endpoint. Use a dedicated serving environment and follow the minimal server recipe:

CUDA_VISIBLE_DEVICES=0,1 SGLANG_DP_SIZE=2 \
bash serve_llada_ui.sh /path/to/hf_ckpt_unpacked

Then replay the packaged mobile, desktop, and web requests:

export SGLANG_BASE_URL=http://127.0.0.1:30000/v1
export SGLANG_MODEL=LLaDA-UI

python3 inference/sglang_client.py --example mobile
python3 inference/sglang_client.py --example desktop
python3 inference/sglang_client.py --example web

Each JSON embeds its screenshot and an evaluated current-image-only multi-turn request. The packaged step-$t>0$ cases use the exact role sequence system, user(""), assistant(previous), user(current): the empty historical user turn is followed by the latest raw <think>...</think><action>...</action> response, while the final user turn contains text first and exactly one image, the current screenshot. Historical screenshots are never replayed. Mobile keeps its task/history headings in the current user caption; desktop and web keep the task in the system prompt and use Current Screenshot: as the current caption. At step 0, omit the empty user and previous assistant messages. Send another compatible request with python3 inference/sglang_client.py --request request.json.

GUI Action Format

For navigation, the model maps the task, current screenshot, and available interaction history to tagged reasoning and a text-serialized action. A typical response has the following form:

<think>Reasoning grounded in the current GUI state.</think>
<action>Click(box=(x, y))</action>

Coordinates are integer-normalized to [0, 999]. The action vocabulary is platform aware:

  • Mobile: Click, DoubleClick, LongPress, Drag, Swipe, Type, LaunchApp, Wait, CallUser, GetScreenshot, device navigation, Answer, and Finished.
  • Desktop: pointing actions, RightClick, Drag, Swipe, Type, Hotkey, Wait, CallUser, and Finished.
  • Web: pointing actions, Drag, directional Scroll, Hover, Type, URL Launch, Hotkey, browser navigation, CallUser, and Finished.
  • Grounding: directly returns [x,y], with [-1,-1] for an infeasible request, rather than returning a navigation action.

Following the dominant training-data format, point-based navigation actions use box=(x,y). Here box denotes one normalized interaction coordinate rather than a rectangular region.

Citation

@techreport{lladaui2026,
  title  = {LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents},
  author = {Zhangxuan Gu and Haoxing Chen and Qi Qin and Yi Xin and
            Kai Gan and Lin Liu and Long Cui and Xiaomei Wang and
            Beitong Zhou and Yunzhu Zhang and Zhengwen Zeng and
            Changlong Gao and Weizhi Chen and Rongchao Zhang and
            Haoyuan Wu and Shuheng Shen and Changhua Meng and Weiqiang Wang and
            Jianguo Li and Zhenzhong Lan},
  year   = {2026},
  url    = {https://huggingface.co/inclusionAI/LLaDA-UI}
}

About

No description, website, or topics provided.

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages