Introducing computer use in Gemini 3.5 Flash

78d ago · US · primary source: deepmind.google

Google has integrated computer-use capabilities directly into its Gemini 3.5 Flash model, a move that lets developers build custom agents that can see, reason, and act across browser, mobile, and desktop environments [1]. The feature was previously available only as a standalone Gemini 2.5 computer-use model but is now a native tool within the main Flash variant [1]. Developers can access the capability through the Gemini API and the Gemini Enterprise Agent Platform [1]. Google says the integration improves performance on long-horizon enterprise automation tasks such as continuous software testing and knowledge work across professional applications [1]. Gemini is a family of multimodal large language models developed by Google DeepMind, succeeding LaMDA and PaLM 2, and was first announced in December 2023 [2][3]. The Flash tier is positioned as a cost-effective, high-throughput variant within the broader Gemini lineup, which also includes Nano, Pro, and Ultra versions [2]. The 1.5 and 3 model generations introduced extended context windows capable of analyzing entire codebases or long-form videos in a single prompt [2]. In a demonstration, 3.5 Flash used its computer-use tool to analyze the Gemini application and return a categorized list of features [1]. It also audited its own documentation for accessibility issues [1]. The release comes as agentic browsers—web browsers that can navigate pages or fill out forms on a user’s behalf—gain traction. Several such browsers, including ChatGPT Atlas, Comet, and Dia, emerged in 2025 [5]. Established browsers like Chrome and Edge have also added AI features, with Chrome integrating the Gemini chatbot for U.S. desktop users [5]. To address prompt-injection risks for agents operating in live environments, Google applied targeted adversarial training to the computer-use feature in 3.5 Flash [1]. The company is also releasing two optional enterprise safeguards: one that requires explicit user confirmation for sensitive or irreversible actions, and another that automatically stops tasks if an indirect prompt injection is detected [1]. Google recommends that developers combine these features with secure sandboxing, human-in-the-loop verification, and strict access controls [1]. A demo environment hosted by Browserbase is available for testing, and reference implementations are provided through the Gemini API and Enterprise Agent Platform [1].

product-launch

Background sources we checked (6)
  • en.wikipedia.org ↗ Gemini (also known as Google Gemini and formerly known as Bard) is a generative artificial intelligence chatbot and virtual assistant developed by Google. It is powered by the family of large language models (LLMs) of the same name, after previously being based on LaMDA and PaLM …
  • en.wikipedia.org ↗ Gemini is a family of multimodal large language models (LLMs) developed by Google DeepMind, and the successor to LaMDA and PaLM 2. Comprising Gemini Pro, Gemini Deep Think, Gemini Flash, and Gemini Flash Lite, it was announced on December 6, 2023. It powers the chatbot of the sam…
  • en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…
  • en.wikipedia.org ↗ An AI browser is a web browser with integrated artificial intelligence capabilities, such as automatically summarizing web page content or answering questions about it. A more specialized type is an agentic browser, based on the concept of agentic AI, which can take actions – suc…
  • en.wikipedia.org ↗ x402 is an open, neutral payment standard for internet-native transactions built on the HTTP protocol. It repurposes the long-unused HTTP 402 "Payment Required" status code to enable peer-to-peer payments directly within HTTP request-response cycles. Payments can be made in suppo…
  • en.wikipedia.org ↗ Virtual worlds are playing an increasingly important role in education, especially in language learning. By March 2007 it was estimated that over 200 universities or academic institutions were involved in Second Life (Cooke-Plagwitz, p. 548). Joe Miller, Linden Lab Vice President…

Sources covering this (5)

Spot something wrong? Report an issue