AI Dev Pulse–2026-09-24

At a glance

  • Cursor shipped Teams/Enterprise Security Review (exploitable bugs on the PR) and Rollouts (PR-to-prod health: verified healthy / regression detected / inconclusive), with trial credits for about 50 (Teams) or 500 (Enterprise) Rollouts changes over 10 days.
  • Same day, Cursor's harness work cut measured user token cost 7% with no quality drop they track: shorter system prompt, dynamic tools, tighter cache breakpoints, compressed file reads.
  • Claude Code v2.1.281: Bedrock gateway assume_role, guardrail: {id, version} on every Bedrock upstream or none, telemetry.resource_attributes, optional "attribution": false, and stricter claude plugin validate.
  • Gemini CLI v0.61.0 hardens indirect prompt injection and sandbox boundaries; it does not patch Plugin4Shell.
  • VS Code v1.139 runs agent sessions inside Dev Containers over SSH, Tunnel, and WSL; Codex v0.156.1 puts GPT-6 Sol and Luna in the picker.

Yesterday set the model picker. Today is the last mile after the agent writes the diff: who reviews it for exploitable bugs, who watches the first environments it lands in, and whether the client you actually run will honor a Bedrock guardrail or a Dev Container toolchain. The Opus 5.5 and GPT-6 Sol rows are already there; the new work is review, rollout, and client policy on one service you already ship.

Treat today as the day you attach a reviewer to the PR, a monitor to the deploy, and a policy file to the agent client, instead of debating which model should be the org default.

Top Stories

Cursor: Security Review on the PR, Rollouts on the deploy

Practical dev impact: Cursor’s 23 Sep product post and changelog ship two bots aimed at the work after code exists. Security Review reads each pull request against the rest of the repo and posts one comment on exploitable issues (injection, auth gaps, secrets, unsafe deserialization, known-vulnerable dependency bumps, insecure defaults), while style and quality stay with Bugbot. Rollouts attaches a monitor to the PR, writes a monitoring plan from the diff before merge, then watches the change through your deploy system and telemetry (Datadog, Grafana, Honeycomb, or equivalent), and per environment it reports verified healthy, regression detected, or inconclusive; it can ping the author, pause a rollout, or revert depending on how you configure it. Feature-flag ramp and release-train awareness are listed as coming soon, not shipping today. Both bots are Teams and Enterprise: enable from the dashboard automations tab, connect source control, and for Rollouts also connect the deploy system and a telemetry provider. Cursor is including usage credits for roughly 50 (Teams) or 500 (Enterprise) Rollouts changes for 10 days. Pick one production service, not the whole org; turn Security Review on that repo this morning, turn Rollouts on only if you can name the deploy pipeline and the latency/error boards it should read, and do not let it auto-revert until the first week of verdicts looks right.

Cursor also cut the token bill 7% without a new model

Practical dev impact: A separate 23 Sep research note (Katz, O’Keefe, Yee) says A/B tests on live users cut token cost 7% with no drop in the quality metrics they track (usage, latency, tool-call errors). The work is harness-side: system prompt trimmed about 66%; built-in tools moved to dynamic context (60% fewer static tool-description tokens); earlier MCP-on-demand load had cut tokens 46.9% on sessions that actually called an MCP tool; cache breakpoints after stable layers cut cold misses 20%; Read tool line numbers only on every tenth line cut cache-read tokens 1.6%. Those are Cursor-measured numbers on Cursor’s distribution, so they are not a reason to skip routing and they are a reason to stop pasting every MCP server and every tool schema into the static prefix of your own agents. Audit the always-on MCP list in Cursor and Claude Code today; if a server is used in fewer than one in five sessions, load it on demand; keep the system prompt, repo map, and tool schemas stable so cache hits survive; and do not “optimize” by deleting the one tool the last-mile reviewer needs.

Claude Code v2.1.281: gateway policy, not another model id

Practical dev impact: Claude Code v2.1.280 (22 Sep) is the Opus 5.5 wire-up; v2.1.281 (23 Sep) is the enterprise control plane. On Claude apps gateway Bedrock upstreams you can now assume_role through STS, including a role in another AWS account, optionally one session per developer, and you can set guardrail: {id, version} on those upstreams with the changelog rule that it is all Bedrock upstreams or none. telemetry.resource_attributes stamps fixed labels on Desktop and /login sessions, and "attribution": false in settings.json hides commit and PR attribution (older CLIs skip a settings file that uses that key, so keep the object form in any file shared across versions). On the laptop side, claude plugin validate now reports .mcp.json entries that would be dropped at load, undeclared ${user_config.*} references, and insecure URLs; MCP URL-mode elicitation can open a browser flow without leaving a stuck waiting dialog; and /insights estimates how many recent permission prompts auto mode could have handled. The Claude Code GitHub Action bumped the bundled CLI to v2.1.281 the same day. If the team hits Bedrock through a gateway, land assume_role plus one guardrail id in the gateway config before you raise the default model; run claude plugin validate on every marketplace plugin already installed; pin the Action to a SHA that includes v2.1.281+; and do not set "attribution": false in a repo settings file that still has to load on v2.1.280 machines unless you keep the object form.

Gemini CLI v0.61.0 hardens the loop. It does not close Plugin4Shell.

Practical dev impact: Stable v0.61.0 posted 23 Sep. Highlights include stopping indirect prompt injection through build-file modifications and untrusted command flags, hardening filesystem boundaries and isolating internal AgentLoop state so spreads do not drop required fields, and keeping explicit versioned Flash model ids. Install remains npm install -g @google/gemini-cli; compare v0.60.0...v0.61.0. Google still tells unpaid and Google One users that Gemini CLI was replaced by Antigravity CLI on 18 Jun 2026. Plugin4Shell status is unchanged from the 17-18 Sep disclosure: Google declined an agent-side pin-verify fix and pointed people at Antigravity, so v0.61.0 is a different class of bug (injection plus sandbox) and you should not treat the stable bump as the Plugin4Shell patch. If you still run Gemini CLI on a paid seat, move the laptop and any CI image to v0.61.0 today and keep marketplace plugins off; if the org already accepted the Antigravity migration, do not keep a v0.60.x binary around “just in case.” Copilot still had no agent-side Plugin4Shell patch at disclosure; GitHub’s public position remains a host-side ban on SHA-like ref names on github.com.

VS Code v1.139 and Codex v0.156.x: put the agent in the real environment

Practical dev impact: VS Code v1.139 (23 Sep) extends Dev Container agent sessions from local folders to projects reached over SSH, Remote Tunnels, and WSL via chat.agentHost.devContainer.enabled (Agents Window). The agent host also stopped opening every conversation database to draw the session list; Microsoft says large lists load and refresh faster (Visual Studio Magazine reported up to 12x in Microsoft’s test), and compact view, empty-group filters, and in-place rename (F2) shipped in the same notes. Codex CLI v0.156.0 (22 Sep) made worktree sessions default, added optional fullscreen /tui, a /usage dashboard, and voice-on-by-default (F8 to toggle); v0.156.1 (23 Sep) is the hotfix that puts gpt-6-sol and gpt-6-luna in the model picker and names Luna in the rate-limit switch prompt. If the service already has a Dev Container, turn chat.agentHost.devContainer.enabled on and run the next agent session there instead of hoping the laptop toolchain matches CI. On Codex, confirm v0.156.1+, pick Sol for the daily lane, and turn voice off on shared machines (/voice or F8). Worktrees-by-default means parallel Codex jobs should no longer share a dirty tree, so check that your disk and git hooks can stand the extra checkouts.

Practical Impact Analysis

The through-line is the same job in three layers, and none of them is another model-routing debate.

Layer one is the diff. An agent that can call Opus 5.5 or GPT-6 Sol still ships SQL injection and leaked secrets unless something that is not the authoring model reads the PR, so Cursor’s Security Review is one productized version of that check while Bugbot remains style, not a substitute for exploitable-bug review.

Layer two is the deploy. Token routing from yesterday does not tell you whether checkout latency in one region is the change you just merged. Rollouts only works if source control, the deploy system, and telemetry are actually connected, and an inconclusive verdict is a missing dashboard, not a green build.

Layer three is the client. Claude Code v2.1.281 is how a Bedrock shop puts STS and a guardrail in front of whatever model id you froze yesterday. Gemini CLI v0.61.0 is how a remaining Gemini CLI seat closes a different hole than Plugin4Shell. VS Code v1.139 is how the agent stops compiling against the laptop’s Node while CI uses the Dev Container. Codex v0.156.1 is how the Sol/Luna price card becomes a picker row instead of a blog post.

Do the three layers on one service today. Do not announce a company-wide bot rollout and a gateway rewrite in the same afternoon.

Tutorial

Last-mile checklist on one repo you already ship. About thirty minutes. No new model ids. Numbered steps first; then one copy-paste block.

  1. Freeze the clients: Claude Code wants v2.1.281+ (Plugin4Shell floor remains v2.1.179); Codex wants v0.156.1+ (Plugin4Shell floor remains v0.146.0); Gemini CLI wants v0.61.0 if you still run it.
  2. From the repo root, run claude plugin validate and fix or delete anything it flags (dropped .mcp.json entries, undeclared ${user_config.*}, insecure URLs) before you enable a new bot.
  3. On Cursor Teams/Enterprise, open https://cursor.com/automations, enable Security Review for a single repository, leave Bugbot on for style, open a throwaway PR with an obvious sink (unsanitized query into SQL, or a hardcoded test secret), confirm one review comment lands, then revert. Enable Rollouts only after you can point it at the deploy system and the Datadog / Grafana / Honeycomb board for that service; do not grant revert on day one. If you are not on Cursor Teams, use the same throwaway PR against whatever security reviewer you already pay for; the point is a bot that is not the author.
  4. In VS Code user or workspace settings, set "chat.agentHost.devContainer.enabled": true, open the remote folder (SSH, Tunnel, or WSL), pick Use Dev Container in the Agents Window, and run one build/test prompt there. If the container cannot build, fix that before you blame the model.
  5. On Codex, confirm v0.156.1+, pick gpt-6-sol for daily implementation (Luna when the rate-limit prompt fires), turn voice off on shared desks and CI images (/voice or F8), and check /usage so the new ids are the ones actually billing.
  6. Bedrock shops only: in the Claude apps gateway Bedrock upstream, set assume_role to the role the developer should wear and set guardrail: {id, version} on every Bedrock upstream or on none. Stamp telemetry.resource_attributes so Desktop and /login sessions show up on the same board as the CLI.
bash Tutorial
#!/usr/bin/env bash
set -euo pipefail

echo "== Claude Code (want v2.1.281+; Plugin4Shell floor v2.1.179) =="
claude --version

echo "== Codex (want v0.156.1+; Plugin4Shell floor v0.146.0) =="
codex --version

echo "== Gemini CLI (want v0.61.0 if you still run it) =="
gemini --version || true

# From the repo root: dropped .mcp.json entries, undeclared ${user_config.*}, insecure URLs
claude plugin validate


... click "Show full code" below to expand
▸ Show full code (28 lines)
#!/usr/bin/env bash
set -euo pipefail

echo "== Claude Code (want v2.1.281+; Plugin4Shell floor v2.1.179) =="
claude --version

echo "== Codex (want v0.156.1+; Plugin4Shell floor v0.146.0) =="
codex --version

echo "== Gemini CLI (want v0.61.0 if you still run it) =="
gemini --version || true

# From the repo root: dropped .mcp.json entries, undeclared ${user_config.*}, insecure URLs
claude plugin validate

# VS Code / workspace setting (Agents Window Dev Container on remote):
# { "chat.agentHost.devContainer.enabled": true }

# Codex picker hygiene (interactive):
# /model  -> gpt-6-sol (daily) ; gpt-6-luna when rate-limit prompt fires
# /voice  -> off on shared desks and CI images
# /usage  -> confirm the new ids are billing

# Bedrock gateway (all Bedrock upstreams or none):
# assume_role + guardrail: {id, version}
# telemetry.resource_attributes for Desktop and /login

echo "Next: enable Security Review on one repo; Rollouts only after deploy+telemetry are named."

Confirm client versions, a clean claude plugin validate, one throwaway-PR review comment, and (if applicable) Dev Container agent host on before you widen the bot rollout.

Recommended AI prompt

Copy this paragraph into ChatGPT, Claude, Gemini, Grok, or whatever you use.

You are a staff engineer helping me wire last-mile review, rollout monitoring, and client policy now that the model picker is already set. Context from 23 Sep 2026: Cursor shipped Security Review and Rollouts for Teams and Enterprise (trial credits about 50 Teams / 500 Enterprise Rollouts changes for 10 days); Cursor also published a 7% harness token cut with no measured quality drop; Claude Code v2.1.281 adds Bedrock gateway assume_role, guardrail on every Bedrock upstream or none, telemetry.resource_attributes, optional attribution false, and stricter claude plugin validate; Gemini CLI v0.61.0 hardens injection and sandbox boundaries but does not patch Plugin4Shell; VS Code v1.139 enables Dev Container agent sessions over SSH, Tunnel, and WSL via chat.agentHost.devContainer.enabled; Codex v0.156.1 puts gpt-6-sol and gpt-6-luna in the picker. I will paste my current Cursor, Claude Code, Codex, and VS Code setup plus one service name and whether we hit Bedrock through a gateway. Return (1) which last-mile bot to enable first on that service, (2) the exact client version floors I should enforce today, (3) whether Bedrock guardrails apply and what to put in the gateway config, and (4) one throwaway-PR test I can finish in about 30 minutes. Do not re-litigate model prices, do not recommend a company-wide bot rollout, and do not invent dashboards or deploy systems I did not name. First reply: the four-item checklist only; wait for my paste before expanding.

Go deeper in Grok

Explore each Top Story in Grok. Links open in a new tab. On phones, the same link may open the Grok app if you have it installed (via your device's normal link handling).

Article: AI Dev Pulse–2026-09-24

Privacy: links open grok.com in your session only. AIDevPulse does not run your prompts through our API.

Leave a Comment