菜单
小墨

小墨

PhoneBuddy: Hybrid Training Framework Achieves 83.2% Success Rate on AndroidWorld

The evolution of large language models has shifted from answering questions to directly operating software interfaces. In the mobile domain, this manifests as AI agents that must complete real-world tasks: booking hotels, filling documents, searching for information in mini-apps. The critical difference lies in understanding that task success is not measured by whether a button is correctly identified, but whether the agent can complete the entire task given the device's current state—account lo

概述

PhoneBuddy addresses a fundamental challenge in mobile agent training: the tension between realism and scalability. Real app environments provide authentic feedback but are expensive, slow, and difficult to automate for evaluation. Simulated environments offer scale and automatic verification but risk learning behaviors that don't transfer to real devices. The research team proposes combining both approaches—using real apps for ground truth feedback while leveraging PhoneWorld, a reconstructed s

Training Architecture

The framework starts all models from the same Qwen3.5-4B backbone with shared action interfaces and SFT initialization using 950,758 action steps from both real and simulated trajectories. After supervised fine-tuning, models diverge into two reinforcement learning branches: Real-only RL and Real+Mock hybrid RL with a 50/50 rollout ratio. The real app environment uses rubric-based model judging—Gemini generates evaluation rubrics and Qwen3.5-122B-A10B scores trajectories against them. PhoneWorld

PhoneBuddy's value lies not in beautiful numbers but in clarifying the most awkward contradiction in mobile agent training: the more realistic the environment, the harder it is to scale; the more scalable the environment, the more it loses fidelity.

“Research Analysis”
积墨 AI 核心产品

积墨 AI 智能体开发平台

快速搭建具备商业价值的 AI 智能体,支持复杂工作流编排、50+ 主流模型接入与私有化部署。

Performance Results

The progression from SFT to Real to Real+Mock shows consistent improvement on single-app and AndroidWorld tasks. On AndroidWorld specifically, success rates climb from 60.3% to 77.2% to 83.2%—demonstrating transferable gains beyond the paper's internal evaluation set. Single-app tasks reach 62.0%, outperforming Gemini 3.1 Pro (50.0%) and GPT-5.4 (50.0%) despite using a much smaller model. Average success rate across all categories reaches 54.8%, representing a 5.0 percentage point improvement ov

Cross-App Limitations and Future Directions

The most revealing finding is the persistent weakness in cross-app tasks, which show no improvement across training stages (22.0% → 20.0% → 18.0%). This highlights that single-app practice doesn't naturally extend to scenarios requiring information transfer between applications—copying content from one app to another, maintaining context across interfaces, or coordinating actions across different apps. Such capabilities would require specialized designs for state tracking, temporary storage, cro

如有侵权,请联系删除。

#AI Agent#Mobile AI#Open Source#Deep Learning
分享文章

相关文章推荐

试用咨询
企业微信二维码

扫码添加企业微信