
SWE-bench Verified moved from 60% to nearly 100% in a single year.
Wow…
A benchmark built to separate frontier labs from everyone else stopped doing that job in roughly the time it takes most enterprises to renew a cloud contract.
Capability is the part of AI that keeps solving itself.
Stat 1: The benchmark that stopped meaning much
That SWE-bench jump comes from Stanford HAI’s AI Index Report, alongside a blunter finding: industry produced 91% of notable AI models in the tracked period, up sharply from prior years. Capability gains have become the easy part of this story.
A near-perfect SWE-bench score leaves open whether that same model can sustain a coherent multi-turn conversation, how often it still gets things wrong once deployed, and whether anyone can actually demonstrate that reliability beyond the benchmark itself.
The report’s own framing lands harder than most vendor decks manage: a field scaling faster than the systems around it can adapt. Somewhere, a marketing team is already turning that sentence into a keynote slide, missing the point on the way to the font choice.
Stat 2: The number that should worry a CISO more than the benchmark score
Veracode’s GenAI Code Security Report tested more than 100 models and found the average security pass rate sitting at 56%, essentially flat against the prior year. Java code failed security tests over 70% of the time.
Models write code that compiles at close to a perfect rate. Writing code that resists an attacker turned out to be a separate skill entirely, one that matters most once that code reaches mission-critical systems. The compiler has zero opinion on your firewall.
The same disconnect shows up in agent evaluations, where models keep failing simple selection traps even as headline scores climb.
Stat 3: The transparency score moving in the opposite direction from capability
Stanford’s Foundation Model Transparency Index fell from 58 to 40 points this year, and 84% of the most notable recent models shipped with the training code omitted. Google, Anthropic, and OpenAI have all stopped disclosing dataset sizes and training durations for their newest releases.
The labs topping the leaderboards are the same labs publishing the least about how they got there.
Stat 4: The money that says nobody is waiting for the safety data
Two figures put a number on how much confidence the market is placing in capability alone, ahead of the trust question catching up.
- Global corporate AI investment hit $581.7 billion, up 130% year over year, with generative AI investment alone reaching $170.9 billion, according to Stanford HAI.
- The top four US hyperscalers pushed combined data center capex toward $600 billion, including Meta, on a path toward $1.7 trillion globally by 2030, per Dell’Oro Group.
Increasingly, that spending follows a “train once, infer forever” logic that is reshaping how budgets get allocated.
Stat 5: The language shift that happened faster than most roadmaps updated
GitHub’s Octoverse report found TypeScript overtook both Python and JavaScript in August 2025 to become the platform’s most used language, the biggest language shift GitHub has recorded in over a decade. Somewhere, a Python maintainer is recalculating a career plan mid-standup.
A five-year language roadmap became a two-year one, the kind of shift AI architects now plan around faster than most individual engineer checklists get updated.
Stat 6: The skill cluster that outran the org chart
Lightcast’s contribution to the Stanford AI Index tracked mentions of the “agentic AI” skill cluster in US job postings and found them growing over 280% in a single year, moving from 0.06% of postings to 0.23% (roughly 90,000 postings).
Hiring managers are asking for a skill set that barely had a name eighteen months ago, while a lot of resumes still list “proficient in Excel” as the differentiator right above the part where they claim they can wrangle an agent swarm.
Stat 7: The workforce number sitting in survey data before it hits headcount
The same Stanford Index found a third of surveyed organizations expect AI to shrink their workforce in the coming year, concentrated in service operations, supply chain, and software engineering.
Large-scale job losses have yet to appear in aggregate employment data, a gap the report treats as a timing issue rather than a reassurance. That distinction matters more than either half of the sentence read alone, and it is worth revisiting at the next earnings season rather than this one.
Stat 8: The adoption line that crossed mainstream before governance caught up
Organizational AI adoption reached 88%, up from 78%, according to the same Stanford report. Almost every organization worth surveying has adopted something, which in plenty of cases means a pilot license nobody has opened since the kickoff meeting.
Adopting something differs meaningfully from capturing enterprise value from it, and differs further still from building the agent experience layer most of these systems still need. That gap is exactly what shows up in the growing list of agentic deployment mistakes enterprises keep repeating.
So, what should builders and buyers do with the gap?
- Treat benchmark scores as a single input among several. SWE-bench near 100% and a security pass rate stuck at 56% describe the same model generation. Procurement built on capability scores alone is reading half the page, a lesson every applied machine learning product manager learns early, usually right after a vendor demo that skipped security entirely.
- Ask vendors what they stopped disclosing, alongside what they shipped. A transparency index falling from 58 to 40 points is a sourcing question every technical buyer should be asking out loud.
- Budget for the security review the Veracode number implies. A 56% pass rate means roughly half of committed AI-generated code needs the scrutiny a team gives human-written code, at minimum.
- Hire for the skill cluster before the job posting catches up. A 280% jump in agentic AI mentions is a leading indicator worth acting on before it turns into a talent shortage, especially while the agents themselves keep tripping over basic “why” questions. Interview accordingly.
Capability was always going to move first. The genuinely hard problem, the one these numbers leave open, is closing the distance between how good these systems have become and how much anyone actually knows about them.
That problem lands on engineers first, but it lands just as hard on the chief AI officers steering strategy from Silicon Valley and the tech leaders doing the same from New York. Progress got its keynote. Trust is still waiting for its turn on stage.
Where the infrastructure money actually has to land
The capex figures above describe the top of the funnel. What happens between a hyperscaler writing a check and a production AI system running is a modernization problem most infrastructure teams are still solving in real time.

AIAI’s report, Bridging the gap from supercomputing to AI factories, pulls direct insight from NVIDIA and WEKA leadership on that exact gap: where HPC design assumptions break under production AI workloads, what modernization looks like once the check clears, and how teams keep existing infrastructure investment intact while making the shift.
Download the report before the next infrastructure budget gets approved on assumptions worth checking first.


