Alphabet Inc.’s Google has begun rolling out Gemini 4 Argon, its latest flagship artificial intelligence model, amid internal debate over its performance, particularly in coding applications. The company introduced Gemini 4 to a select group of cybersecurity partners on Wednesday, with plans to expand availability to paying subscribers after further trials.

Google reported that Gemini 4 achieved top scores on various benchmark tests, including outperforming OpenAI’s Astra model on a security-related evaluation. However, some employees familiar with the project expressed concerns that these benchmarks do not fully reflect the model’s real-world capabilities. According to insiders, Gemini 4 struggles with certain coding tasks, especially front-end design, an area critical for applications and user interfaces. These critiques highlight a disparity between impressive test results and practical utility.

Despite skepticism voiced by some staff, Google leadership remains optimistic. Koray Kavukcuoglu, head of Google DeepMind, recently reiterated his confidence in the model’s development and the team behind it. Opinions within Google appear divided, with some employees viewing Gemini 4 as trailing competitors like Anthropic’s Fable and OpenAI’s Astra, while others believe the new model matches or surpasses rival technologies.

Google faces significant pressure for Gemini 4 to succeed, given its extensive integration across many of the company’s key products, including Search, Maps, Gmail, and Chrome, each with over a billion users. The model’s performance could influence Google’s ability to maintain dominance in AI-driven search and software services while competing with firms that are increasingly focusing on user-facing AI products such as coding agents.

The company’s recent AI development timeline has seen setbacks. A planned release of Gemini 3.5 Pro, intended for June, was abandoned. Such delays incur substantial costs, with analysts estimating up to $400 million in training expenses for large AI models. Additional challenges include the model’s large size, which may increase operational costs.

Experts note that Google’s focus on optimizing benchmark scores—a practice termed “benchmaxxing”—might contribute to issues with practical application. This tendency can prioritize high test performance over usability or versatility in real-life coding scenarios. Some insiders argue that while Gemini 4 excels at processing diverse inputs, such as extracting metadata from video, and shows strengths in cybersecurity and safety, its coding capabilities remain uneven.

Google’s internal environment reflects broader industry pressures, with some key AI researchers departing recently and leadership changes at DeepMind. Demis Hassabis shifted to a chairman role in August, passing day-to-day operations to Kavukcuoglu.

Google maintains that Gemini 4 is designed to handle complex, extended tasks across software engineering, finance, legal work, and cybersecurity. The model can generate long outputs, reportedly up to 1 million tokens, or about 750,000 words, per instance. Meanwhile, competitors continue rapid advancement. Meta Platforms Inc. recently released Muse, an AI agent capable of completing everyday tasks and quickly rising in popularity.

As Google rolls out Gemini 4, its long-term success will depend on translating benchmark achievements into robust, practical performance to maintain leadership in the competitive AI landscape.