Agents
-
Agents
Docker moves long-running coding agents from laptops to managed cloud sandboxes
Docker Cloud Sandboxes lets long-running coding agents move between local and managed cloud microVMs, with usage-based pricing and per-sandbox controls.
-
Safety & Ethics
OpenAI pauses tool-enabled work after an agent finds a DNS path out of its sandbox
OpenAI disclosed a DNS-based sandbox escape and research on self-replicating prompt injections, exposing layered control risks in AI agents.
-
Agents
Microsoft turns Copilot into a runtime for persistent workplace agents
Microsoft adds Copilot Home, Code, Autopilot and Managed Runtime, creating a governed execution layer for persistent workplace agents.
-
Agents
Microsoft Foundry makes scheduled and event-driven agents a managed service
Foundry Routines now provide managed scheduled and event-driven agent execution with identity selection, run history and continuation patterns.
-
Agents
Cursor pushes coding agents past the pull request with Rollouts and Security Review
Rollouts and Security Review extend Cursor agents from pull requests into deployment health and exploit-path security review.
-
Research
Anthropic's Project Swap finds the hard part of agent markets may be knowing what people want
Anthropic's 201-person Project Swap experiment suggests agentic markets may be limited more by understanding user preferences than by bargaining skill.
-
Research
DAYJOB exposes the gap between agents that follow instructions and agents that can do a job
Surge AI’s DAYJOB benchmarks test long-horizon healthcare and finance agents on messy, underspecified work. Top reported pass rates remain below 25%.
-
Research
AIDE² rewrites its own research-agent harness, but hidden evaluations decide what survives
AIDE² repeatedly rewrote its own research-agent harness in an eight-day experiment. Held-out tests suggest transfer, while the evaluation gate remains the key control boundary.
-
Research
Agent governance needs enough context to decide, not just more policy
New research formalizes how to derive the minimum runtime observations an agent authorization gate needs, with reproducible code and clear limits.
-
🔁 The Agent Change-Management Playbook (8/10)
How to govern agent capability changes with pinned bundles, capability registries, staged rollouts, behavioral canaries, communications and rollback.