At a glance
- GitHub's Project HydraFusion research preview routes Copilot CLI tasks through Single, Cascade, or Critique multi-model workflows.
- Claude Code v2.1.261 adds
/skill-doctorand 128K tool-output caps; v2.1.263 is the install pin after the Sep 2 to Sep 6 wave. - Cursor Self-Hosted Machines move Cloud Agent tool execution onto your workers while inference stays in Cursor's cloud.
- GitHub Actions adds a runner-deprecation API, a
vulnerability-alertstoken permission, and reusable-workflowjob.workflow_*identity fields.
September 7 is a Labor Day Monday after a week of model launches, and the useful signal sits in how you run the agents you already have. That means which workflow gets chosen, how unattended sessions stay clean, and where tool execution actually runs, rather than another flagship drop. This brief stays on those verified agent-ops moves instead of padding a holiday morning.
Treat today as an agent-ops day. Try HydraFusion on one well-scoped first-turn coding task, pin Claude Code to v2.1.263 and run /skill-doctor, and decide whether Cloud Agent work needs Cursor-managed VMs or your own workers.
Top Stories
Project HydraFusion brings multi-model orchestration to Copilot CLI
Practical dev impact: "Which model?" is no longer the only cost lever, because HydraFusion picks a Single, Cascade, or Critique workflow per request, and GitHub's offline numbers show large cost cuts with quality that matches Opus 5 on only one of three benches. GitHub published the research preview on September 4. HydraFusion builds an execution plan across providers, then chooses one of three patterns: Single, where one model solves it; Cascade, where an efficient draft meets a quality gate that can escalate; or Critique, where a draft is reviewed by a tool-less critic from another model family and then revised once. Versus Claude Opus 5, TerminalBench 2.1 improved verified quality by 4.9 points at 67% lower estimated cost, DeepSWE was 1.5 points below at 36% lower cost, and CheckpointBench was 0.1 points below at 65% lower cost. It is available on all Copilot plans in Copilot CLI as a research preview. GitHub says first-turn, single-prompt coding tasks are the best place to start, and multi-turn is still in progress. Usage bills at each underlying model's standard token rates across every workflow leg.
Claude Code v2.1.261 ships /skill-doctor; pin v2.1.263 for the full Sep wave
Practical dev impact: If you run headless or CI agents, upgrade past Friday's v2.1.260 pin, then audit skill context cost before you widen unattended permissions. Release notes for v2.1.261 (September 4) add /skill-doctor to show which loaded skills go unused and what they cost in context, raise inline Bash and background-task output caps to 128K characters via bashOutputMaxChars and taskOutputMaxChars, add an Organization policy line to /status and claude doctor, and add --append-subagent-system-prompt-file for large subagent prompts. Version v2.1.263 (September 6) is the reliability pin, and there is no v2.1.262 in the registry. Read the wave with v2.1.259 as well: managedMcpServers covers org-provided HTTP/SSE MCP servers, and --permission-prompts none makes unattended print-mode sessions deny instead of hanging on prompts. Prefer @anthropic-ai/claude-code@2.1.263 over assuming stable already moved.
Cursor Self-Hosted Machines put Cloud Agent tool execution on your infra
Practical dev impact: If network, GPU, Mac or iOS, or custom image constraints block Cursor-managed Cloud Agents, move only the worker, not the agent loop. Cursor's September 2 post and docs describe Self-Hosted Machines: install the Cursor CLI, run agent worker start, and open a long-lived outbound HTTPS connection so tool calls (edits, terminal, computer use, local MCP) run on your machine while inference and planning stay in Cursor's cloud. There are two shapes. My Machines is for personal workers, and Team Pools (Enterprise) uses a controller that scales capacity from a spawn script. Partner sandbox paths include AWS Lambda, Cloudflare, Coder, Daytona, E2B, Modal, Namespace, and Vercel. Linux workers now support computer use alongside Mac. Cursor-hosted VMs remain the default, and Self-Hosted is for cases managed cloud cannot meet. Workers need outbound HTTPS to api2.cursor.sh, api2direct.cursor.sh, and (unless you block artifacts) cloud-agent-artifacts.s3.us-east-1.amazonaws.com.
GitHub Actions Early September updates harden runners, tokens, and reusable workflows
Practical dev impact: Plan runner upgrades with the new deprecation API, and stop granting broad token scopes when a workflow only needs Dependabot alert reads. The September 3 changelog adds GET /actions/runners/deprecations/{version} (repo, org, or enterprise) returning runner_version, runtime_deprecates_at, and registration_deprecates_at. Workflows can grant GITHUB_TOKEN a read-only vulnerability-alerts permission (read or none) instead of wider scopes. Reusable workflows get job.workflow_ref, job.workflow_sha, job.workflow_repository, and job.workflow_file_path so a called workflow can identify its own source file at runtime (not available on GitHub Enterprise Server). Pair this with the same week's agent CI work, because least-privilege tokens and known runner expiry dates matter more once agents write PRs unattended.
Practical Impact Analysis
Today's through-line is the control plane for agent work: which models run, where tools execute, and how unattended sessions fail closed. HydraFusion is GitHub's bet that the next gain is workflow selection, not another single default model, so read the table before you trust the marketing. Cost fell on all three offline benches versus Opus 5, but verified quality beat Opus 5 only on TerminalBench 2.1. DeepSWE and CheckpointBench traded a small quality gap for large estimated savings. For an eng lead that means you should pilot HydraFusion on first-turn, well-scoped tasks with a cost dashboard open, and keep a pinned Opus or Fable baseline for repo-level multi-file work until multi-turn orchestration lands.
Claude Code's September 2 to September 6 wave is the unattended counterpart. --permission-prompts none turns hangs into denials you can tune, managedMcpServers keeps MCP config out of every runner image, and /skill-doctor finally prices skill bloat in context tokens. If you pinned v2.1.260 on Saturday for /diff and the parentheses permission fix, move to v2.1.263 today so you also get skill audits and the larger output caps.
Cursor Self-Hosted Machines answer a different constraint: code and secrets that cannot leave your network, or builds that need Macs, GPUs, or custom images. Inference still leaves your network, so treat that as an explicit trade, not air-gapped autonomy. Prefer managed Cloud Agents plus private connectivity when that is enough, and stand up My Machines or Team Pools only when the docs' "who should use this" list matches your blockers.
Actions' Early September trio is the CI floor under all of the above. Know when a runner version loses registration and runtime support, grant Dependabot alert reads without elevating the whole token, and let reusable workflows prove which file actually defined the job. If you only do three things this morning, enable HydraFusion for one single-prompt coding task in Copilot CLI, upgrade Claude Code to v2.1.263 and run /skill-doctor on your heaviest project, and sketch whether Cloud Agent execution belongs on Cursor-managed VMs or a self-hosted pool.
Tutorial
Enable HydraFusion in Copilot CLI, then pin Claude Code v2.1.263 and audit skills. Use a substantial, well-scoped first-turn coding task (GitHub's preview guidance). Expect HydraFusion in the model picker after experimental mode is on, then confirm Claude Code is on the September 6 pin before you trust /skill-doctor.
- In Copilot CLI, run
/update, then/experimental, then/modeland select HydraFusion. - Hand it one well-scoped first-turn coding task and watch cost across every workflow leg.
- Pin
@anthropic-ai/claude-code@2.1.263and confirmclaude --versionreports v2.1.263. - In a real repo with skills loaded, run
/skill-doctorand prune only skills you agree are dead weight. For unattended CI later, add--permission-prompts noneonly after permission rules are tuned.
Recommended AI prompt
Copy this paragraph into ChatGPT, Claude, Gemini, Grok, or whatever you use.
You are my staff engineer for AI agent ops on 2026-09-07. GitHub Project HydraFusion (Copilot CLI research preview) chooses Single, Cascade, or Critique workflows per request, with large estimated cost cuts versus Claude Opus 5 and quality that beat Opus 5 on only one of three offline benches. Claude Code v2.1.261 adds /skill-doctor and 128K tool-output caps; pin v2.1.263 after the v2.1.259 managed MCP and --permission-prompts none wave. Cursor Self-Hosted Machines move Cloud Agent tool execution to your workers while inference stays in Cursor cloud. GitHub Actions Early September adds a runner-deprecation API, vulnerability-alerts token permission, and reusable-workflow identity fields. Ask which harnesses we use and whether code must stay on our network. Then produce (1) a one-page HydraFusion pilot plan with a pinned single-model baseline, (2) a Claude Code v2.1.263 upgrade checklist focused on /skill-doctor and unattended --permission-prompts none, and (3) a go or no-go note for Cursor Self-Hosted Machines versus managed Cloud Agents plus a 15-minute Actions checklist. Keep it concrete and copy-paste ready.