C:\CHANGELOG> type v1-9-1-rollout-report-aug-17.md
v1.9.1 · released · 5 min read · by

Rollout Report: 750 tokens a second into a very small funnel

Weekly Rollout Report. The theme is throughput. Two labs made the machine dramatically faster, Gartner says your accounting software is growing an agent whether you asked or not, and one guy spent two months making a flagship phone worse. There's a lesson in the order of those.

Frontier intelligence at 750 tokens a second

OpenAI previewed Ultrafast, an API service tier running GPT-5.6 Sol at up to 750 output tokens per second — up to 14x its standard speed, same model, not a distilled cousin. The speed comes from Cerebras wafer-scale silicon, which keeps the weights on-chip instead of streaming them from external memory. On Humanity's Last Exam it finished all 2,500 questions in 11 hours and 11 minutes. On GDP-Val, a benchmark for economically valuable knowledge work, OpenAI reports a 5.6x end-to-end speedup with no quality degradation. Limited preview, no pricing announced.

Operator read: latency is a UX feature, not a benchmark. For the batch work that pays my bills — takeoff extraction, invoice reconciliation, certified payroll checks — four seconds instead of forty changes nothing, because the PM opens it Thursday either way. Where 750 tok/s genuinely matters is anything a human is standing there waiting on: voice on a jobsite, an estimator poking at a drawing, live support. Real category, newly viable. Not most of my backlog.

Claude Code stopped asking

On August 14, Anthropic made auto mode the default for Pro, Max and Team, routing each action through a separate classifier that blocks anything escalating beyond your request, touching infrastructure it doesn't recognize as yours, or driven by hostile content the model just read. Enterprise and the raw API stayed opt-in.

The justification is the number I can't stop thinking about: 97% of Claude Code permission prompts get approved. In a study of 1,053 paid testers, humans caught dangerous commands 13.6% of the time; the classifier caught 89%. So the prompt was never a control. It was a speed bump with a rubber stamp bolted to it, and Anthropic finally said so out loud. Same finding as the human review gate, from the other side — a gate only works if the person standing at it can realistically fail something. Note what did not change: deny rules still block outright and the classifier can't override them. Judgment got automated; the hard boundary stayed in config, which is the only place a boundary has ever counted.

The agents are shipping pre-installed

Gartner now expects 40% of enterprise applications to ship with task-specific AI agents built in by the end of 2026, up from under 5% a year earlier. On the infrastructure side, Oracle said OCI is among the first clouds to support NVIDIA's Nemotron 3.5 Lightning, an open model built specifically for always-on agents.

This reframes the whole small-business conversation. Nobody is going to sit down and "decide to adopt agents." Your accounting package, your CRM and your PM tool are each going to grow one on a Tuesday, in a release note. The governance question stops being should we build this and becomes which vendor's agent is already inside my general ledger, what can it write, and who reads its output. Go inventory that while it's still a spreadsheet and not an incident.

Contech: the money keeps going to paperwork

Eight construction tech startups raised in the week ending August 3, per Bricks & Bytes: two rounds into AI touching drawings and infrastructure design, three into trades workforce and payroll compliance, three into materials, modular and logistics. SoftBank is also reportedly in talks to buy Swiss robotics firm Gravis Robotics for north of $500m.

Three of eight going to workforce and prevailing-wage paperwork is the entire industry in one line. The robot arm gets the headline; the money follows certified payroll, because that's what's actually bleeding. The sexiest problem in construction is a bricklaying robot. The most expensive one is a wage determination nobody can audit.

And the odd one: two months to make a Galaxy S9+ worse

NTDEV, the developer behind tiny10 and tiny11, got full Windows 10 Enterprise LTSC booting on a Samsung Galaxy S9+ after two months of nights and a great deal of AI-assisted trial and error. It runs. It also has reduced available RAM, no sound, and no network.

I love this without reservation, and it's the cleanest demo-versus-deployment parable I've seen all year. The impossible part worked. The two things a user notices in the first ten seconds did not. That's every pilot I've been handed to clean up: the deliverable was never the boot screen.

Deployed next week: inventorying every agent feature my vendors have quietly shipped into Coen's stack — what it can write, and whether anybody signed off. If Gartner's 40% is even half right, the audit I do in August is a lot cheaper than the one I do in January.

Everybody bought a bigger hose this week. Nobody bought a bigger funnel.

— Cole Ciprari · Business Systems Architect · Worcester, MA
my résumé is an operating system → ciprari.ai · linkedin.com/in/coleos · cole@ciprari.ai
▚▞ GET THE NEXT RELEASE
New releases Monday, Wednesday and Friday, plus the Sunday Rollout Report — the week's AI and tech news, summarized by a human with production access. No spam. Unsubscribe by emailing a mildly disappointed cole@ciprari.ai.
PHOSPHOR