When AI starts to take actions: GUI-Agent Smartphone Hacking

We analyze an AI-native smartphone where a language model acts as an execution layer, perceiving screen content and autonomously performing user interactions (e.g., clicks, swipes) across apps—effectively granting a GUI agent system-level control.

We present a security analysis of this architecture and showcase its security concerns. Despite built-in safeguards around sensitive operations and personal data, we demonstrate that the system can be systematically steered into performing attacker-controlled actions, enabling access to private data (e.g., memory, messages, media) and even participation in multi-stage attack chains.

We further study the evolving defence mechanisms deployed on the cloud side and observe a continuous attacker–defender dynamic and a clear tradeoff between security and usability. Our findings include a prompt-based exploitation technique that remains effective against recent mitigations (as of Feb. 2026), as well as insights into cloud-side logic obtained after achieving root access on the device.

This work highlights the security implications of treating LLMs as execution layers and frames prompt injection as a practical, persistent exploitation primitive in real-world smartphone systems.

 

About the Presenter: Li Shi

Li Shi is a security researcher at DARKNAVY, specializing in fuzzing and binary exploitation, discovered 100+ vulnerabilities in real-world products prior to the AI era. Currently focused on AI security and applying AI to security research. Previous work has been published at deepsec.cc, Usenix Security, NDSS  and IEEE S&P.

Previous
Previous

Cyber Reasoning Systems for the Next Generation

Next
Next

AI Security at Scale: Memory Safety Across the ML Inference Stack