News archive
October 8, 2026
Claude Haiku 5.5 brings cheaper agent work to a high-volume model
Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens up to 100k context. Its lower price makes it a serious first pass for repetitive agent work.
$0.10/$0.50 input/output price per million tokens72.4% OSWorld score75% lower average cost than Haiku 4.5
GPT-6 ships in ChatGPT with interactive UI responses
GPT-6 puts interactive output beside text in ChatGPT and reaches 1.2B weekly users. The useful product question is which actions become faster when the answer is already a working interface.
1.2B weekly ChatGPT users44% faster start to web answers than GPT-5.6 InstantOct 7/8 paid/free rollout dates
OpenAI adds a Decisions API for typed choices and scores
The Decisions API is aimed at small, repeated judgments rather than open-ended replies. OpenAI says it is 10x faster than the Responses API and charges $0.10 per million input tokens.
10x claimed speed over Responses API$0.10 input price per million tokens1 model in public beta
Mistral Large 4 pairs a one-million-token window with a low launch price
Mistral Large 4 offers a 1M-token context, multimodal input, and launch pricing of $0.68 per million input tokens. An open-weight release is planned for the end of October.
1.05T/52B total/active parameters1M token context window$0.68/$2.09 input/output price per million tokens
EmbeddingGemma 2 puts multimodal search inside a phone-sized model
EmbeddingGemma 2 can run with about 191MB of active RAM for text-only use and 567MB for full multimodal use. It targets local search and intent routing without fine-tuning.
740M parameters191MB text-only active RAM on Pixel 11 Pro567MB full multimodal active RAM
ChatGPT adds audio upload and meeting-note workflows for paid users
Paid ChatGPT users can upload audio up to 512MB, ask questions about a recording, and turn a meeting or lecture into notes. Free-plan audio support is not included.
512MB maximum audio upload size8 listed audio formatsOct 6 release-notes date
Google Playground turns text prompts into shareable games
Playground lets users prompt a game, revise it, and share it through a gallery. Google says multiplayer and leaderboards are available for selected genres, with access tied to rollout and subscription level.
18+ launch age requirementUS initial launch market3 create, share, publish steps
Docker Agent v1.149.0 adds GitHub skills and a Decisions evaluator
Docker Agent v1.149.0 turns agent setup into YAML, supports multiple model providers and MCP servers, and now loads skills from GitHub repositories. Docker Desktop 4.63+ includes it.
v1.149.0 latest release on Oct 74.63+ Docker Desktop version5+ listed model-provider families
Google opens SynthID Detector for public watermark checks
SynthID Detector can check supported AI media from Google and partners, but a negative result is not proof that media is human-made. Watermark coverage remains the key limit.
180B Google-stated watermarked images and videos240,000 years Google-stated watermarked audio1M+ daily verification requests in listed products
OpenAI publishes formal math progress without releasing the underlying model
OpenAI's public math repository lists 722 manuscripts in 372 families and says each result used about three hours of ChatGPT Pro thinking compute on average. The model itself is not released.
722 manuscripts in the public catalogue372 result families3 hours average ChatGPT Pro thinking compute per result
Musubi releases PolicyLM-1.7B for fast custom moderation decisions
PolicyLM-1.7B is designed for local policy checks, with a reported 35ms median on an L4 and 19-language support. Musubi reports 83% accuracy on its custom-policy benchmark.
1.7B parameters35ms reported median per short message on L419 languages listed by Musubi
Lambda plans a $4B raise as AI compute demand resets its scale
A reported $4B Lambda raise would value the company at $14.5B before the round. The reported $50B backlog is dominated by a large Anthropic commitment, not verified customer cash.
$4B reported planned raise$14.5B reported pre-money valuation$50B reported September backlog
OpenAI and Ironclad turn contract workflows into computer-use training tasks
OpenAI says GPT-6 Astra scored 32% higher than GPT-5.6 Sol on Ironclad's contract tasks while taking 48% less time per attempt. The result is an evaluation claim, not a general automation guarantee.
32% reported score gain over GPT-5.6 Sol48% reported lower time per attempt
Atlassian connects GPT-6 agents to Rovo's Teamwork Graph
Atlassian says Rovo can use the Teamwork Graph to assess launch readiness, blockers, and missed milestones. More than 3,000 Atlassian developers reportedly use Codex in daily development work.
3,000+ Atlassian developers using Codex3 named work sources: Jira, Confluence, discussions
Anthropic expands cyber access with Project Glasswing partners
Project Glasswing adds three access tiers for advanced cyber work and a partner group spanning cloud, security, hardware, and open-source organizations. Access is for qualifying security professionals.
3 Cyber Verification access tiers11 named Project Glasswing partners
FemtoAI opens developer access to a low-power AI chip stack
FemtoAI is opening access to an SPU kit, compiler, and Sparsity Studio. The company claims 10x lower power and 10x smaller memory, but builders still need device-level tests.
10x claimed lower power10x claimed smaller memory use1 SPU evaluation kit
GitHub rebuilds Git infrastructure for agent-scale repository traffic
GitHub reports 473.3B monthly Git events, 7.38B commits in September, and 3.26B Actions runs. Agent-scale development is now an infrastructure problem, not only a model problem.
473.3B monthly Git events in Aug 20267.38B September commits3.26B September Actions runs
Caterpillar and CoreWeave shorten physical-AI training loops
Caterpillar reportedly manages 18PB of federated machine data and wants simulation and reinforcement learning to cut physical-AI feedback to hours. The report describes an industrial workflow, not a ready-made product.
18PB reported Caterpillar federated dataterabytes reported data from one machine per dayhours target feedback time instead of weeks or months

















