LaunchedEditorial Listing

ARTEMIS

Google · ARTEMIS: Google's Open-Source AI Agent for Android Test Automation

Open ARTEMIS

ARTEMIS is an open-source (Apache 2.0) agent from Google that turns natural-language instructions into automation on real Android phones or emulators. It ships a web console, a CLI, a Python SDK, and an MCP server so coding agents such as Claude Code, Codex, and Antigravity can drive a device. It suits mobile developers and QA engineers who want AI-driven device testing and bug reproduction without writing UI scripts first.

PricingFree
Setupmedium
Runs onSelf-hosted
APIYes
Open sourceYes
DocsYes
CategoryCoding
Mobile TestingAndroidTest AutomationMCPOpen SourceBug ReproductionQA TestingPython SDK

Best for

Android developers and mobile QA engineers who want an AI agent, driven from their coding assistant or test suite, to operate real devices for testing and bug reproduction

Not ideal for

iOS teams, testers who want a hosted device lab with no local setup, and teams that cannot send device screenshots to a cloud model provider

Who it's for

Android developers, mobile QA and test engineers, and teams using AI coding assistants who need to test on real devices

Capabilities

  • Runs natural-language tasks on real Android phones (with USB debugging) or emulators over ADB, including cross-app workflows
  • Flash profile: a single-model observe-and-act loop at about 3 to 5 seconds per step for routine UI tasks
  • Pro profile: Planner, Operator, and read-only Checker agents with a living task plan, pre-execution checks on individual actions, and verification levels from off to strict
  • Element locating through accessibility hierarchies and OCR, with coordinate and visual fallbacks for Canvas, Compose, and Flutter UIs
  • MCP server with tools to run and manage tasks, capture screenshots or UI hierarchy, inspect traces, and diagnose the environment
  • One-command MCP and rules installation for Antigravity, Cursor, Claude Code, Codex, Windsurf, VS Code, Cline, Roo Code, and OpenClaw
  • Local web console with device connection wizard, live screen mirroring, prompt sandbox, and execution replays
  • `artemis run` CLI for scripted tests, exploratory stability checks, and AndroidWorld benchmark runs
  • Python SDK client with typed (Pydantic) results for assertions in pytest and CI pipelines, talking to an ARTEMIS host over HTTP
  • Collects Logcat output, screenshots, and session recordings for bug reproduction and reports
  • Completion notifications through desktop toasts, webhooks (Slack, Discord, CI), or custom script hooks

Limitations

  • Android only. iOS support is listed on the roadmap, not available
  • You supply your own model API keys and pay the provider. The default configuration uses Gemini models
  • According to the default config, the object detector used for visual coordinate grounding must run on a Gemini Embodied Reasoning model (it warns that non-ER models 'will fail spatial coordinate detection'), so setups without Gemini access lose that fallback for custom UIs
  • Installs an accessibility helper app on the device the first time a task runs (it can be disabled in favor of UIAutomator2)
  • Pro mode is much slower per step (about 15 to 40 seconds) than Flash, and Flash has no plan, checkpoint verification, or ADB shell
  • Requires a local setup with ADB, Python 3.12+, and a connected device or emulator. There is no hosted version
  • The 99%+ AndroidWorld result is reported by the project itself
  • Young project (repository created August 2026) with no tagged releases yet and frequent changes

Use cases

  • Asking Claude Code or Codex to build an APK, install it on a connected phone, walk through the login flow, and return screenshots
  • Reproducing a reported bug on a real device and collecting Logcat output and a recording
  • Running long exploratory or stability sessions with the Pro profile and a written report
  • Adding natural-language device checks to a pytest suite through the Python SDK
  • Exploring a live app first so a coding agent writes UI test code based on verified interaction paths

Our take

ARTEMIS fills a gap most coding agents leave open: they can write and build mobile code but cannot tap through the app on a phone. Exposing a real device through MCP means an assistant can build, install, and exercise a flow, then report back with screenshots and logs. The split between a fast Flash loop and a slower Pro workflow with a separate Checker lets you trade speed for verification per task. Expect some setup work: you need ADB, a device, and model keys, visual coordinate grounding depends on a Gemini Embodied Reasoning model, and the project is only about seven weeks old. Android teams already using AI coding assistants have the most to gain.

Who should use it

Android developers who use Claude Code, Codex, Antigravity, or Cursor and want them to verify changes on a real device, QA engineers automating exploratory and regression checks, and teams that need to reproduce device-specific bugs with logs.

Who should skip it

iOS-only teams, testers who want a managed cloud device lab, and organizations that cannot run local tooling with ADB access or send screenshots to a model provider.

Strengths

  • Free and Apache 2.0, running on your own machine and devices
  • Works on physical phones as well as emulators
  • MCP server lets coding agents test the app they just built on a real device
  • Pro profile adds planning, per-action checks, and a separate Checker for verification
  • Web console, CLI, and Python SDK cover interactive, scripted, and CI use

Weaknesses

  • Android only for now
  • Visual coordinate grounding depends on a Gemini Embodied Reasoning model
  • Local setup with ADB and model keys required
  • Benchmark claims are self-reported and the project is new

ARTEMIS pricing

Open source

Free

  • Apache 2.0
  • Runs locally with your own Android devices or emulators
  • Bring your own model API keys

Note: ARTEMIS is free. Costs come from the model providers you configure (Google Gemini by default), which scale with the profile you choose and the length of each task.

Technical specs

Modalities

Screenshots, UI hierarchy and OCR, Screen recordings, Text

Where ARTEMIS excels

Verifying a change on a real phone

Through MCP, a coding agent can build and install an APK, run through the changed flow on a connected device, and return screenshots, instead of the developer testing by hand.

Reproducing device-specific bugs

ARTEMIS can follow a bug report's steps on real hardware and capture Logcat output and a session recording for the fix.

Natural-language checks in CI

The Python SDK client returns typed results you can assert on, so natural-language device checks can run inside pytest-based pipelines against an ARTEMIS host.

ARTEMIS vs. competitors

ARTEMIS vs. Momentic Mo

Momentic Mo is a hosted, paid QA agent that bug-bashes web apps and uploaded builds on hosted iOS simulators and Android emulators. ARTEMIS is free, open source, and runs locally against your own Android devices, including physical phones, but has no iOS support or hosted infrastructure.

ARTEMIS vs. Browser Use

Browser Use lets agents operate web browsers through a Python library and cloud browsers. ARTEMIS operates Android devices through ADB and the phone's own interface, for mobile apps rather than websites.

Frequently asked questions

What is Google ARTEMIS?

ARTEMIS is an open-source agent in Google's GitHub organization that takes natural-language instructions and carries them out on a real Android phone or emulator, for testing, bug reproduction, and cross-app automation.

Is ARTEMIS open source and free?

Yes. It is licensed under Apache 2.0 at github.com/google/artemis. The software is free, but you need API keys for the language models it uses and pay those providers directly.

Does ARTEMIS work with Claude Code or Codex?

Yes. ARTEMIS includes an MCP server, and `artemis mcp --install` can configure it for Claude Code, Codex, Antigravity, Cursor, Windsurf, VS Code, Cline, Roo Code, and OpenClaw, along with a testing rules file for the assistant.

Which models does ARTEMIS support?

The environment template accepts keys for Google Gemini, OpenAI, Anthropic, OpenRouter, and xAI. The default configuration uses Gemini models, and the object detector used for visual coordinate grounding requires a Gemini Embodied Reasoning model.

Does ARTEMIS support iOS?

Not yet. iOS devices and simulators are listed on the roadmap. Today ARTEMIS works with Android devices and emulators over ADB.

What is the difference between Flash and Pro in ARTEMIS?

Flash is a fast single-model loop at about 3 to 5 seconds per step, suited to routine UI tasks, without a plan or checkpoint verification. Pro is a multi-agent workflow at about 15 to 40 seconds per step, with a Planner, an Operator, pre-execution checks on individual actions, and a Checker that verifies the result.

Integrations & fit

AntigravityClaude CodeCodexCursorWindsurfVS CodeClineRoo CodeOpenClawADBscrcpypytestGemini APIOpenAI APIAnthropic APIOpenRouterxAISlack and Discord webhooks
Good fit forSolo / individual, Startup / small team, Enterprise
Pricing modelFree· No cost to start
See pricing on ARTEMIS →

Alternatives to consider

About ARTEMIS

ARTEMIS connects to an Android device or emulator over ADB and works through the phone's interface the way a person would: reading the screen, tapping, typing, and moving between apps. It finds elements through accessibility hierarchies and OCR first, then falls back to coordinates and visual models for custom Canvas, Compose, and Flutter interfaces. There are two execution profiles. Flash is a fast observe-and-act loop, typically 3 to 5 seconds per step, for routine tasks, with no task plan, checkpoint verification, or ADB shell. Pro is a multi-agent graph at roughly 15 to 40 seconds per step: a Planner keeps a Markdown plan with checkpoints, an Operator executes it with the full toolset (including ADB diagnostics and video analysis), each individual action passes a pre-execution check against the live UI, and a read-only Checker verifies checkpoints and reviews the result against the original goal. You can use it four ways: a local web console with live screen mirroring and replays, the `artemis run` CLI, a lightweight Python SDK client that calls an ARTEMIS host over HTTP from pytest or CI, or an MCP server with tools such as `mobile_run_task` and `mobile_diagnose`. The setup script (macOS, Linux, or Windows) installs ADB, scrcpy, FFmpeg, and Python dependencies and can register the MCP server and a testing rules file with Antigravity, Cursor, Claude Code, Codex, Windsurf, VS Code, Cline, Roo Code, and OpenClaw. That lets a coding agent build an APK, install it, exercise a flow on a real phone, and return screenshots and Logcat output. You supply model API keys for Google, OpenAI, Anthropic, OpenRouter, or xAI, but the default configuration uses Gemini models, and its comments say the object detector that handles visual coordinate grounding must run on a Gemini Embodied Reasoning model. The project reports a 99%+ completion rate on Google Research's AndroidWorld benchmark. The main limits: it is Android-only for now (iOS is on the roadmap), it installs an accessibility helper app on the device, and it is a young repository created in August 2026 with no tagged releases yet, so expect fast change.

Are you the founder? Claim this listing →