Insights

AI in December 2024: The Reasoning Era Begins (and the Money Follows)

December 2024's AI news—OpenAI's o3, Gemini 2.0, Sora, Databricks' $10B round, and Google's Willow chip—and what each means for builders.

December closed out 2024 with two stories running in parallel. On one track, the labs raced through an advent-calendar season of launches and pushed past chatbots into models that actually reason and act. On the other, capital kept pouring into the infrastructure layer that makes any of it useful in production. For operators, the signal isn’t “AI got better”—it’s that the gap between a demo and a deployed system narrowed again, and the cost of waiting went up.

OpenAI’s o3 Reframed What “Smart” Means

OpenAI ended its “12 Days of OpenAI” run on December 20 by announcing o3, the successor to its o1 reasoning model. The numbers were genuinely unusual: 96.7% on the 2024 AIME math exam, 87.7% on graduate-level science questions, and 25.2% on EpochAI’s Frontier Math benchmark where no prior model had cleared 2%. It also posted a breakout score on the ARC-AGI test designed specifically to resist memorization.

Why it matters for builders: Reasoning models are slower and more expensive per call, so they’re not a drop-in replacement for everything. The opportunity is selective: route hard, multi-step problems—complex extraction, planning, code review—to a reasoning model, and keep cheaper models for the routine 80%. The teams that win in 2025 will architect for tiered model use, not pick one model and call it done.

Google Shipped Gemini 2.0 and Aimed at Agents

On December 11, Google introduced Gemini 2.0 Flash, framing it explicitly around “the agentic era.” It outperformed the prior 1.5 Pro on key benchmarks at roughly twice the speed, with native multimodal input across text, images, audio, and video—plus early image and text-to-speech generation.

Why it matters for operators: Speed and multimodality at a low price point change what’s economical to build. Workflows that were too slow or too costly to run on every document, call, or ticket become viable. Before you commit to one vendor, benchmark Gemini 2.0 against your actual tasks—the “best model” depends entirely on your workload and latency budget.

Sora Turbo Put Video Generation in Real Hands

On December 9, OpenAI released Sora Turbo publicly to ChatGPT Plus and Pro subscribers. Text-to-video left the research-demo phase and became a product people could actually use.

Why it matters: For most B2B teams this isn’t core infrastructure, but it reshapes marketing, prototyping, and content economics. The practical move is to treat it as a cost lever for creative work—not a reason to staff up. Know what it can do so you’re not paying agency rates for what a tool now handles.

Databricks Raised $10B—The Infrastructure Bet Is On

On December 17, Databricks announced a $10 billion Series J at a $62 billion valuation, led by Thrive Capital. The capital is earmarked for AI products, acquisitions, and go-to-market expansion, on the back of 60%+ year-over-year growth.

Why it matters: The smart money is concentrating on the data-and-deployment layer, not just the models. That tracks with what we see in practice—most AI value stalls on plumbing: getting the model clean access to your systems and data. The headline models are commoditizing fast; the durable advantage is in how well your business is wired to use them.

Google’s Willow Reminded Everyone the Frontier Is Wider Than LLMs

On December 9, Google unveiled Willow, a 105-qubit quantum chip that completed a benchmark in under five minutes that would take a classical supercomputer an effectively unimaginable span—and, more importantly, showed error rates dropping as qubits scaled.

Why it matters: Nothing here changes your roadmap this year or next. But it’s a useful reminder that “AI strategy” and “compute strategy” aren’t the same thing, and that the long arc has more than one frontier. File it under awareness, not action.

December’s lesson is that the tooling is moving faster than most roadmaps. If you’re weighing where reasoning models, multimodal pipelines, or a tiered model architecture fit your 2025 plan—and what to skip—let’s talk through it.