Google tests Gemini 4 update as coding competition grows

Google employees are reportedly testing Gemini 4's Carbon version, but its public release and coding improvements remain unconfirmed.

Lauren Collins ·

Google tests Gemini 4 update as coding competition grows

Google employees are reportedly testing Gemini 4's Carbon version, but its public release and coding improvements remain unconfirmed.

The internal testing comes as Google prepares to introduce Argon publicly, following its September announcement of the Gemini 4 generation. Employee assessments suggest Carbon may offer stronger coding capabilities, although those impressions do not establish how it performs against competing products.

Documents and screenshots described in an October 9 report indicate that Carbon recently became accessible through Jetski, Google's internal environment for coding. Google declined to comment, leaving both the model's status and its relationship to the planned Argon release unresolved.

Argon's name obscures internal versions

The development pipeline involves more than a straightforward progression from one product name to another. An internal document described in the report lists Gemini 4 versions called Argon, Barium and Carbon, but those development labels do not necessarily match the names customers encounter.

That distinction matters for the coming launch: the document identifies Barium-B as the version selected for public introduction under the Argon name. Carbon may be a later checkpoint, meaning an updated version, rather than a separately marketed product; the available information does not settle that question.

Staff comparisons stop short of benchmarks

One employee with access to the internal model said Carbon "feels like Opus 5.5" when used for coding, comparing it with an Anthropic model. The employee also cautioned that further testing was necessary, making the comment an initial assessment rather than a demonstrated performance result.

Earlier Argon versions received broadly favorable employee reactions, according to conversations and internal messages described in the report. One employee nevertheless judged some coding capabilities closer to Anthropic's older Claude Opus 5, illustrating how impressions can differ across versions and tasks.

Neither comparison comes with published scores, a common set of tasks or an independent evaluation in the supplied material. The employees were not identified, and the report did not explain their anonymity; their access provides a view of internal testing, not independent confirmation of Google's competitive position.

Cybersecurity testing precedes wider access

Google's own account of Argon emphasizes more than software development. In its September announcement, the company claimed leading performance on selected programming and professional-task benchmarks, including work involving finance and law; those claims remain separate from employees' descriptions of Carbon.

The company also emphasized defensive cybersecurity and said Fairwind Program partners would receive access ahead of a broader release. Google described that program as a way to examine cyber vulnerabilities in new models, placing an additional testing stage between its announcement and wider availability.

The company separately announced a workplace-focused Gemini agent during the week of the report. That product announcement adds context to the model testing, but it does not establish that Carbon will power the agent or become available to customers.

If Carbon reaches users and reproducible tests support the employee assessments, the update would give developers firmer grounds to compare Google's coding tools with Anthropic's offerings. If it remains internal, the relevant product test will instead be the publicly released Argon version; no Carbon launch date or final public name has been established.

More stories

Latest news