At a glance
- Artificial Analysis Index v4.3 ties Claude Fable 5.1 and GPT-6 Astra after harder terminal and workflow evals.
- Anthropic's Claude agents delivered a Lean-checked formalization of Fermat's Last Theorem in 11 days.
- LangChain moves MCP into the main package so tool servers stay sessionless and elicitation pauses as interrupts.
- Hugging Face open-sources funes, a local-first memory layer coding agents can recall across sessions.
September 8 is a post-Labor Day Tuesday with few brand-new flagship drops, so the useful work is how these four moves fit together. You still have to score agents against a harder eval, give multi-day jobs a harness that can finish, keep tool servers sessionless, and make sure yesterday's reasoning is still there when the next session starts.
Treat today as an eval-and-continuity day. Re-read your model pick against Index v4.3, where the headline score is tied and the cost is not, then study the FLT harness if you run multi-day agents, migrate MCP clients onto MCPAdapter, and give one coding agent a funes memory before you start another cold session.
Top Stories
Artificial Analysis Intelligence Index v4.3 hardens agentic coding and workflow evals
Practical dev impact: If you are still choosing models from last week's v4.2 scorecard, that number is already stale. v4.3 upgrades Terminal-Bench to v4.0 and swaps τ³-Banking for AutomationBench-AA, and Fable 5.1 and Astra both land at 53, with Astra about 57% cheaper per Index task. Artificial Analysis published the update on September 7 as another interim step toward Index v5. Category weights stay Agents 30%, Coding 20%, General 30%, and Scientific Reasoning 20%, while private-test weight rises from 40% to 45%. On Terminal-Bench v4.0 (66 tasks, pass@1 averaged over three runs), GPT-6 Astra (max) scores 59.1% versus 52.0% for Claude Fable 5.1 (max with fallback) and 49.0% for Claude Opus 5 (max). On AutomationBench-AA (657 held-out Zapier workflows), Astra (max) scores 68.5% on objectives completed with no guardrail violation, and 41.6% of those workflows are fully completed clean, while Fable 5.1 (max with fallback) fully completes 32.1%. Average cost per Intelligence Index task is about $3.26 for Astra (max) versus $7.63 for Fable 5.1 (max with fallback). Fable still leads on AA-Briefcase and SciCode, while Astra leads on Terminal-Bench v4.0 and AutomationBench-AA Score.
Claude agents deliver the first Lean-checked formalization of Fermat's Last Theorem
Practical dev impact: Long-horizon agent work fails when it lives in a single forever chat, because the agents lose project state. Anthropic's September 4 research post reports the first complete computer-checked FLT proof in Lean, written largely autonomously, after Prove2Me plus a Claude Code multi-agent harness finished the job in 11 days. Early uncoordinated runs stalled. The campaign produced about 13 million lines of Lean and about 29,500 intermediate theorems in the final proof (about 30,300 proved along the way), checked against Lean's three standard axioms, with the statement matched to Mathlib's FLT. Switching to Prove2Me, a DAG of theorem statements with a statement and proof file split and natural-language search for reuse, unblocked parallel work. Anthropic cites roughly six billion output tokens from an internal research model roughly comparable to Claude Fable 5.1. The proof is on GitHub at anthropics/fermats-last-theorem. Kevin Buzzard reviewed and called the autoformalization extraordinary. This is verification of an existing proof path, not a novel FLT argument.
LangChain folds MCP into the main package with elicitation and catalog caching
Practical dev impact: If you still pin langchain-mcp-adapters or MultiServerMCPClient, move to MCPAdapter under langchain.mcp and wire human-in-the-loop MCP asks as LangGraph interrupts instead of sticky sessions. LangChain's September 3 post tracks the July MCP rewrite to a stateless core. Install with pip install "langchain[mcp]" (requires langchain[mcp]>=1.4.0, Python first, TypeScript soon, API still beta). The stack sits on FastMCP for transports, auth, negotiation, and caching across old and new protocol eras. Elicitation, meaning mid-tool questions such as confirm-delete, surfaces as interrupt(), and you resume with Command(resume={"responses": {key: answer}}) once a checkpointer holds the paused run. Servers can advertise tool-list TTLs so clients stop refetching catalogs every run (cache=True, cache_mode="use"). ClientGroup keeps per-server auth and prefixes tool names, so billing_search stays distinct from docs_search.
Hugging Face ships funes, durable memory your coding agents own
Practical dev impact: Pasting last week's design rationale into a fresh Claude Code or Codex session is how that decision disappears. Index the traces you already have and let the next agent recall them. Hugging Face published funes on September 3 as a local-first memory layer for Claude Code, Codex, pi, and Hermes. One funes add <agent> builds a Lance index from existing session logs, installs turn hooks, and exposes recall and get tools. Embedding and reranking run on-device with a pinned local model, and results return original turn text with provenance, not distilled summaries. Bind funes add codex org/funes-memory to publish to a Hub dataset you own (private by default) with credential redaction before upload. funes ask answers one grounded question without changing agent setup. Hugging Face reports recall was cheaper than written handoffs on their handoff-vs-recall bench (about 8x and 4x on two tasks) and more reliable than compaction when summaries flattened key findings.
Practical Impact Analysis
Today's through-line is continuity under harder tests. Index v4.3 raises the bar on terminal coding and business workflow automation, then shows Fable 5.1 and Astra tied at 53, with Astra much cheaper per Index task. If your eval harness still quotes Terminal-Bench v2.1 or older banking benchmarks, your leaderboard is already stale, so rebuild the smoke suite against Terminal-Bench v4.0 style terminal work and an automation bench that fails a task on any guardrail trip, then decide default models with cost-per-task open beside the score.
The FLT result is the multi-day harness lesson. Uncoordinated agents burned cycles until Prove2Me gave them an immutable statement graph, separable proofs, and reusable search. If your swarm loses the plot after a weekend, copy that shape: a shared DAG of goals, statement versus proof separation, and a verifier (Lean, tests, or typecheck) that is the only merge gate. Do not treat the 13M-line artifact as Mathlib-ready. Anthropic notes it is likely much longer than it needs to be.
LangChain's MCP move is the production counterpart of that same idea. Stateless MCP plus interrupt elicitation means you can scale tool servers without sticky sessions, and you can pause a destructive tool for a human without inventing a one-off callback. Migrate adapters this week if you still pin langchain-mcp-adapters.
funes closes the loop on session amnesia across hosts and harnesses. Pair it with the FLT lesson, because durable shared state beats heroic context windows. If you only do three things this morning, re-rank your default coding model on Index v4.3 cost and Terminal-Bench v4.0, install langchain[mcp] and point MCPAdapter at one HTTPS server with a checkpointer, and run funes add on the agent you use most so tomorrow's session is not a stranger.
Tutorial
Install funes for Claude Code, then ask one grounded question from local memory. Use this when you already have Claude Code session history on disk and want recall without standing up a hosted memory service. Hub publish is optional. Do not paste secrets into a public memory dataset.
- Install the funes binary (local embeddings; no Hub account required for local recall).
- Index existing Claude Code sessions and install
recall/gethooks withfunes add claude. - Ask one grounded question with
funes ask. - Start a normal Claude Code task that depends on a prior decision and confirm the agent calls
recallwith provenance. Ifasksays the passages do not support an answer, rephrase or wait for indexing to catch older sessions.
Recommended AI prompt
Copy this paragraph into ChatGPT, Claude, Gemini, Grok, or whatever you use.
You are my staff engineer for AI agent evaluation and continuity on 2026-09-08. Artificial Analysis Index v4.3 ties Claude Fable 5.1 and GPT-6 Astra at 53 after harder Terminal-Bench v4.0 and AutomationBench-AA evals, with Astra cheaper per Index task. Anthropic's Claude agents finished a Lean-checked FLT formalization in 11 days via Prove2Me plus a multi-agent harness. LangChain moved MCP into langchain.mcp with MCPAdapter, elicitation-as-interrupt, and tool-list caching. Hugging Face funes indexes Claude Code / Codex session history locally so the next agent can recall it. Ask which models, harnesses, and MCP servers we run. Then produce (1) a one-page Index v4.3 re-rank of our defaults with cost-per-task, (2) a Prove2Me-style DAG checklist for any multi-day agent job, (3) an MCPAdapter migration checklist, and (4) a 15-minute funes rollout for our primary coding agent. Keep it concrete and copy-paste ready.