Insights

AI in May 2025: The Month Agentic Coding and AI Hardware Went All-In

May 2025 brought Claude 4, Google I/O's Gemini 2.5 push, OpenAI's $6.5B Jony Ive deal, and an algorithm-discovery breakthrough. What it means for builders.

May 2025 was the month the frontier labs stopped talking about chatbots and started building for two things at once: code that writes itself and hardware you don’t have to look at. Two model releases reset the bar for coding agents, a $6.5B acquisition signaled where the interface is headed, and a research result quietly broke a 56-year-old mathematical record. For operators, the signal underneath the noise is consistent: the value is moving from “the model can answer” to “the model can do the work.” Here’s what actually happened and how to read it.

Claude Opus 4 and Sonnet 4 raised the coding bar

On May 22, Anthropic released Claude Opus 4 and Claude Sonnet 4, both hybrid reasoning models that can alternate between thinking and tool use, including web search, mid-task. Anthropic positioned Opus 4 as the strongest coding model available and built both to sustain long-running, multi-step agent workflows rather than one-shot answers.

Why it matters for builders: The differentiator is no longer raw IQ on a benchmark — it’s how long a model holds context and stays coherent across a real task. If you tried to wire up coding agents a year ago and got brittle, drifting results, that constraint just loosened. This is the month to re-test workflows you’d previously written off as not-ready.

Google I/O went all-in on agents and video

At I/O on May 20, Google shipped a dense slate: Gemini 2.5 Pro and Flash updates with a new “Deep Think” reasoning mode, AI Mode in Search, and Veo 3 — a video model that generates native audio, dialogue, and sound effects alongside the footage. It also launched Flow, a filmmaking tool stitching its image, video, and language models together.

Why it matters for operators: AI Mode in Search is the one to watch. If a meaningful share of your traffic comes from Google, the SERP is becoming an answer engine, not a list of links. That’s a demand-gen question, not a model question — and it deserves a place on your roadmap now, not after the traffic dips.

OpenAI made two big bets: coding and hardware

OpenAI had an aggressive month. On May 6 it agreed to acquire AI coding startup Windsurf for ~$3B (a deal that later unraveled), a day after rival Cursor closed $900M at a $9B valuation. Then on May 22, OpenAI announced it was acquiring io, Jony Ive’s hardware startup, in a deal valued at ~$6.5B to build a screen-free, context-aware AI device.

Why it matters for builders: The coding-tool land grab tells you where the labs see near-term enterprise money — agentic software development. The Ive deal is a longer bet that the next interface isn’t an app at all. You don’t need to act on the hardware story yet, but the coding story is immediate: assume your engineering org’s tooling will look different in twelve months.

AlphaEvolve broke a record that stood since 1969

Quietly, mid-month, Google DeepMind unveiled AlphaEvolve — a Gemini-powered agent that evolves and verifies its own algorithms. It found a way to multiply 4×4 complex matrices in 48 scalar multiplications, beating Strassen’s 1969 record of 49, and improved on the best known solutions for roughly 20% of 50-plus open math problems it was pointed at.

Why it matters: This is a preview of AI as a discovery engine, not just a drafting assistant. The pattern — generate, automatically verify, keep what wins — is exactly the loop that makes agents trustworthy on real work. Where you can define a hard success check, you can let a model search far harder than a human would.

The takeaway

The through-line for May is that “agentic” stopped being a buzzword and started being a product category. None of this means rebuilding your stack. It means knowing which of these shifts touches your roadmap — coding velocity, search-driven demand, or a discovery problem you’d parked as unsolvable — and which are still someone else’s problem. That triage is the hard part, and it’s worth getting right before you spend a dollar building.

If you’re weighing where agentic AI actually fits your roadmap, let’s talk — we’ll help you separate the moves worth making now from the ones worth watching.