Gemma 4 Brings Open Models to Browser and Mobile Devices
Paige Bailey of Google DeepMind argues that open models are becoming practical to run on ordinary hardware, including in a browser or on a phone, without sending data to a remote API. Her example is Gemma 4, a family ranging from 2 billion to 31 billion parameters that organizations can download, adapt and fine-tune; which version is useful depends on the device and its constraints.

The browser demo makes local inference concrete
Paige Bailey used a browser demo to make the practical case for open models: a model can answer a prompt without sending the request to a remote API. In a web interface, Gemma 4 generated an emoji table comparing the Harry Potter books by humor and excitement, with reading recommendations and a reference to Harry Potter and the Methods of Rationality. The response arrived almost instantly. Bailey questioned the ranking; the demonstration’s point was where the model ran: entirely on the user’s machine.
The demo loaded Gemma in the browser using WebAssembly and Transformers.js. Bailey described the environment as sandboxed and said no data was sent elsewhere. As she put it, “This is not using an API, it's just something that's running locally in the browser with Transformers.js, and the Gemma 4 model.”
Gemma 4 is Google DeepMind’s latest open model family. Bailey listed versions with 2 billion, 4 billion, 12 billion, 26 billion and 31 billion parameters. The 26B model is a mixture of experts; the 31B model is dense. They are Apache 2 licensed, she said, allowing organizations to download and use them in company projects, fine-tune them and adapt them for their own purposes.
Model size determines which hardware is within reach
Bailey framed model size as an infrastructure decision. She said the 26B and 31B models perform beyond what their size might suggest, including compared with models an order of magnitude larger. Their smaller footprint can reduce GPU requirements and avoid distributed inference; the stated aim is to run them on a single commodity GPU.
Smaller variants extend the range of possible deployments. Bailey said some can run on Jetson Nano devices, models of 12B parameters and below can run locally on a laptop, and the 2B model can fit on mobile devices. Quantized checkpoints reduce the footprint further. The fast browser demo used Gemma’s QAT checkpoints, Bailey said; the 2B version can be less than a gigabyte.
Bailey also showed a chart comparing VRAM and storage requirements across Gemma 4 sizes and quantizations, including BF16 and 4-bit Q4_0. Its labels distinguish mobile and mobile text-only options. The comparison makes deployment requirements a matter of both model size and checkpoint format: the model has to fit the hardware available for the intended use.
On-device AI includes image, audio and task skills
Bailey described Google AI Edge Gallery as another way to run models locally. Available for Android and iOS, the app lets users download models and try image description, audio transcription and function calling. Bailey said Gemma supports more than 140 languages.
The gallery also includes skills for tasks such as making games, writing haikus, querying weather and scheduling calendar events. Bailey said the app and locally installed models are free to use. Inference can run on a device accelerator instead of falling back to the CPU; she said Pixel 10 and higher-end mobile devices could already run Gemma locally this way.
Research applications and managed agents extend the picture
Bailey connected Google DeepMind’s work to its stated mission to build AI responsibly and for humanity’s benefit. She named AlphaFold, Med-Gemma, robotics, AI for science, mathematics and frontier research as areas where the organization applies its work, and invited researchers in AI for science to ask how AI might accelerate their work.
She also described a computer-use API and managed agents. Users can give agents a higher-level task in natural language, which they execute in a Linux workstation sandbox. The environment can be extended with skills and dependencies. Bailey presented these tools among the company’s recent releases, alongside open models and on-device applications.
