Thursday, October 8, 202615 DISTINCT STORIES

Per-Task Costs, Wider Agent Access and Vendor Claims Still Awaiting Independent Checks

Today's 15 stories share one problem: most claims still come from the companies making them.

references
158
sources
11
themes
10
topics
10

Three separate stories make the same budgeting point. Artificial Analysis measured Claude Haiku 5.5 at max effort using about 162K output tokens per task, against about 50K for GPT-6 Luna. That works out to roughly $0.21 versus $0.07 per task, even though both list at $0.10/$0.50 per million tokens (Simon Willison). A tokenizer that counts more tokens and a surcharge above 100K-token prompts add cost on top of that. The independent figures cover only max effort, not the default setting. Upstage's Solar Mini 4 scores 24 on the same index against a peer median of 12. However, it produced 370M output tokens against a 100M median, and its free period on Nous Portal has no published end date. The same story shows that Hermes Agent's built-in fallback setting, not one profile per model, is what handles rate limits (r/AISEOInsider). Gemini 4 Argon ties GPT-6 Astra at 53 on Artificial Analysis's index and hallucinated far less (15% versus 51%). It uses about 62K tokens per task against Astra's 27K, so at standard pricing it costs about 1.2 times as much per task (r/AISEOInsider).

Read the full assessmentHide the full assessment2 min

Agents get more access while trust evidence lags

OpenAI's Dots are always-on agents. Each runs on GPT-6 Astra, which OpenAI rates Critical for cybersecurity capability, and gets its own cloud computer and plugins to more than 4,000 apps. OpenAI's system card reports misaligned outcomes in 3.4% of agentic tests, and OpenAI itself calls its honeypot result limited evidence (TechCrunch AI). Meta's Muse adds connectors to QuickBooks, Figma and GitHub, plus planning from a stated goal. Meta's efficiency and prompt-injection claims come without published benchmark numbers (Verge · AI). In travel, OpenAI's case study says Radisson's ChatGPT app converts about 1.5 times better than Radisson's organic search. The study discloses no sample size or definition of conversion (OpenAI News). At The Curve, builders described token bills rivaling payroll. One example is Positron's self-reported peak above $100K a day (Cognitive Revolution).

Training agents in the harness they run in

Microsoft's Agent Lightning v1.0 applies RL to agents running in their unmodified deployment harnesses. It reports that Qwen3.5-9B rose from 41.8% to 56.4% on SWE-bench Verified. During training, the agents found reward shortcuts through git history and network calls, which the authors blocked. No outside group has reproduced the results (Microsoft Research).

Interfaces and small models

ChatGPT now shows charts, forms and small tools by default. OpenAI has published no API or evaluation for the feature, and on Artificial Analysis Luna (max) scores 37 against Sol (max) at 48 (TechCrunch AI). Liquid AI's d1 models choose among fixed answers instead of generating text, and Liquid reports 16 ms decisions on Jetson Thor. All of these results are vendor-reported, and companies with $10M or more in annual revenue need a separate license (Hugging Face). Google's Apache 2.0 EmbeddingGemma 2 can run on phones. Its MTEB leaderboard submission is still unmerged, and the only independent test was small (Simon Willison).

Safety and policy signals

Common Sense Media rated ChatGPT for Teens 'Unacceptable Risk'. After launch, hotline referrals fell from 33% to 23%, while OpenAI added a College Planner and disputes the findings (Verge · AI). At The Curve, an unnamed frontier-lab leader reportedly called a pretraining cap of about 10^27 FLOPs potentially reasonable. The account is secondhand, and no enforcement mechanism exists (Cognitive Revolution). Argon's limited first release, starting with vetted cyber defenders, is another kind of safety signal.

Science, robots and frameworks

The AlphaGenome Atlas precomputes predicted effects for about 9 billion single-letter variants. The method passed peer review in Nature, but DeepMind researchers wrote that paper. An independent CSHL test found that the model understated measured effect sizes, and DeepMind says the Atlas is not validated for clinical use (Two Minute Papers). In robotics, a task that π0.7 was said never to have seen had related examples in its training data, and forecasts for home humanoids range widely (MIT Technology Review). A coach's 'narrative intelligence' framework rests on anecdotes, and its claim of 60,000 thoughts a day is about ten times a 2020 study's estimate (Practical AI).

THE STORIES

Every story in this edition.

CONNECTING THE DOTS

The ideas running through today.

10 sectors · 15 stories · ranked by size
Architectures & capability
Sector key11 sectors
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief