Google · ARTEMIS: Google's Open-Source AI Agent for Android Test Automation
ARTEMIS is an open-source (Apache 2.0) agent from Google that turns natural-language instructions into automation on real Android phones or emulators. It ships a web console, a CLI, a Python SDK, and an MCP server so coding agents such as Claude Code, Codex, and Antigravity can drive a device. It suits mobile developers and QA engineers who want AI-driven device testing and bug reproduction without writing UI scripts first.
Best for
Android developers and mobile QA engineers who want an AI agent, driven from their coding assistant or test suite, to operate real devices for testing and bug reproduction
Not ideal for
iOS teams, testers who want a hosted device lab with no local setup, and teams that cannot send device screenshots to a cloud model provider
Who it's for
Android developers, mobile QA and test engineers, and teams using AI coding assistants who need to test on real devices
ARTEMIS fills a gap most coding agents leave open: they can write and build mobile code but cannot tap through the app on a phone. Exposing a real device through MCP means an assistant can build, install, and exercise a flow, then report back with screenshots and logs. The split between a fast Flash loop and a slower Pro workflow with a separate Checker lets you trade speed for verification per task. Expect some setup work: you need ADB, a device, and model keys, visual coordinate grounding depends on a Gemini Embodied Reasoning model, and the project is only about seven weeks old. Android teams already using AI coding assistants have the most to gain.
Who should use it
Android developers who use Claude Code, Codex, Antigravity, or Cursor and want them to verify changes on a real device, QA engineers automating exploratory and regression checks, and teams that need to reproduce device-specific bugs with logs.
Who should skip it
iOS-only teams, testers who want a managed cloud device lab, and organizations that cannot run local tooling with ADB access or send screenshots to a model provider.
Open source
Free
Note: ARTEMIS is free. Costs come from the model providers you configure (Google Gemini by default), which scale with the profile you choose and the length of each task.
Modalities
Screenshots, UI hierarchy and OCR, Screen recordings, Text
Verifying a change on a real phone
Through MCP, a coding agent can build and install an APK, run through the changed flow on a connected device, and return screenshots, instead of the developer testing by hand.
Reproducing device-specific bugs
ARTEMIS can follow a bug report's steps on real hardware and capture Logcat output and a session recording for the fix.
Natural-language checks in CI
The Python SDK client returns typed results you can assert on, so natural-language device checks can run inside pytest-based pipelines against an ARTEMIS host.
ARTEMIS vs. Momentic Mo
Momentic Mo is a hosted, paid QA agent that bug-bashes web apps and uploaded builds on hosted iOS simulators and Android emulators. ARTEMIS is free, open source, and runs locally against your own Android devices, including physical phones, but has no iOS support or hosted infrastructure.
ARTEMIS vs. Browser Use
Browser Use lets agents operate web browsers through a Python library and cloud browsers. ARTEMIS operates Android devices through ADB and the phone's own interface, for mobile apps rather than websites.
What is Google ARTEMIS?
ARTEMIS is an open-source agent in Google's GitHub organization that takes natural-language instructions and carries them out on a real Android phone or emulator, for testing, bug reproduction, and cross-app automation.
Is ARTEMIS open source and free?
Yes. It is licensed under Apache 2.0 at github.com/google/artemis. The software is free, but you need API keys for the language models it uses and pay those providers directly.
Does ARTEMIS work with Claude Code or Codex?
Yes. ARTEMIS includes an MCP server, and `artemis mcp --install` can configure it for Claude Code, Codex, Antigravity, Cursor, Windsurf, VS Code, Cline, Roo Code, and OpenClaw, along with a testing rules file for the assistant.
Which models does ARTEMIS support?
The environment template accepts keys for Google Gemini, OpenAI, Anthropic, OpenRouter, and xAI. The default configuration uses Gemini models, and the object detector used for visual coordinate grounding requires a Gemini Embodied Reasoning model.
Does ARTEMIS support iOS?
Not yet. iOS devices and simulators are listed on the roadmap. Today ARTEMIS works with Android devices and emulators over ADB.
What is the difference between Flash and Pro in ARTEMIS?
Flash is a fast single-model loop at about 3 to 5 seconds per step, suited to routine UI tasks, without a plan or checkpoint verification. Pro is a multi-agent workflow at about 15 to 40 seconds per step, with a Planner, an Operator, pre-execution checks on individual actions, and a Checker that verifies the result.

Momentic
Product and engineering teams that ship frequently, often with AI-assisted development, and want exploratory and regression QA on web, iOS, and Android without writing or maintaining test scripts
Paid
Browser Use
Python developers and AI teams who need an agent to operate real websites, either by embedding the open-source library in their own stack or by renting hosted stealth browsers and agents on usage-based pricing
FreemiumARTEMIS connects to an Android device or emulator over ADB and works through the phone's interface the way a person would: reading the screen, tapping, typing, and moving between apps. It finds elements through accessibility hierarchies and OCR first, then falls back to coordinates and visual models for custom Canvas, Compose, and Flutter interfaces. There are two execution profiles. Flash is a fast observe-and-act loop, typically 3 to 5 seconds per step, for routine tasks, with no task plan, checkpoint verification, or ADB shell. Pro is a multi-agent graph at roughly 15 to 40 seconds per step: a Planner keeps a Markdown plan with checkpoints, an Operator executes it with the full toolset (including ADB diagnostics and video analysis), each individual action passes a pre-execution check against the live UI, and a read-only Checker verifies checkpoints and reviews the result against the original goal. You can use it four ways: a local web console with live screen mirroring and replays, the `artemis run` CLI, a lightweight Python SDK client that calls an ARTEMIS host over HTTP from pytest or CI, or an MCP server with tools such as `mobile_run_task` and `mobile_diagnose`. The setup script (macOS, Linux, or Windows) installs ADB, scrcpy, FFmpeg, and Python dependencies and can register the MCP server and a testing rules file with Antigravity, Cursor, Claude Code, Codex, Windsurf, VS Code, Cline, Roo Code, and OpenClaw. That lets a coding agent build an APK, install it, exercise a flow on a real phone, and return screenshots and Logcat output. You supply model API keys for Google, OpenAI, Anthropic, OpenRouter, or xAI, but the default configuration uses Gemini models, and its comments say the object detector that handles visual coordinate grounding must run on a Gemini Embodied Reasoning model. The project reports a 99%+ completion rate on Google Research's AndroidWorld benchmark. The main limits: it is Android-only for now (iOS is on the roadmap), it installs an accessibility helper app on the device, and it is a young repository created in August 2026 with no tagged releases yet, so expect fast change.
Are you the founder? Claim this listing →