Over a recent 6-day sprint, a collaborative effort between over 100 AI agents and human developers from Google’s Gemma team and Hugging Face achieved a remarkable fivefold increase in inference speed for the Gemma 4 open models, reaching up to 491.8 transactions per second (TPS) on a single NVIDIA A10G GPU. This initiative reflects a growing trend in AI collaboration, where humans and AI work together intensively to overcome specific technical challenges, particularly in optimizing open model development. However, the fastest result led to a decline in model quality in other areas, showcasing the complexities involved in such optimizations.

@googlegemma: Google Gemma is a family of open-weight AI models developed by Google DeepMind. The Gemma team led a collaborative sprint with Hugging Face to optimize inference performance on Gemma 4. This initiative combined human oversight with AI agent contributions to explore efficiency gains.
Google Gemma: Google Gemma is a family of open-weight AI models developed by Google DeepMind. The Gemma team led a collaborative sprint with Hugging Face to optimize inference performance on Gemma 4. This initiative combined human oversight with AI agent contributions to explore efficiency gains.
Hugging Face: Hugging Face operates a leading platform for hosting, sharing, and collaborating on AI models and datasets. It partnered with the Google Gemma team to run a six-day challenge focused on accelerating inference. The platform enabled pooled resources and self-coordinated efforts among participants.

AI Collaboration: Humans and AI agents are increasingly combined in short, focused sprints to tackle specific technical challenges in model optimization.
Open Model Development: Google advances its Gemma series through partnerships that leverage community platforms for rapid experimentation.