Agents
-
Research
GuardedAct tests a safer boundary for AI-driven incident remediation
GuardedAct reports lower collateral damage when LLM-generated repairs pass through sandbox simulation and an independent execution gate.
-
Agents
Atlassian turns Jira backlogs into governed coding-agent loops
Jira Agent loops can continuously dispatch eligible backlog work to coding agents, while context, standards, review and human merge controls define the boundaries.
-
New Models
GPT-Live-1 brings full-duplex voice to the API, with reasoning delegated behind it
OpenAI releases GPT-Live-1 in the API with full-duplex voice and delegated reasoning, making cancellation, permissions and task-level cost key design concerns.
-
Research
DeFiFlowBench shows why structurally valid agent workflows can still execute unsafe trades
DeFiFlowBench finds that valid-looking DeFi agent workflows can still permit unsafe trades and argues for execution-level policy controls outside the generator.
-
Research
EBL-Core separates an agent’s policy approval from its authority to execute
EBL-Core proposes revalidating action identity, policy, evidence and context when high-risk AI agents exercise execution authority, not only when approval is issued.
-
Research
Agent study finds lost authorization constraints, not context compaction, drive control failures
A Tencent Zhuque Lab study finds agent control failures rise when context management drops authorization rules, while constraint-preserving compaction remains safe in its benchmark.
-
Agents
OpenAI turns the Codex agent harness into a managed API
OpenAI launches Agents API in public beta, separating managed Codex orchestration from selectable execution environments and changing the agent infrastructure boundary.
-
Agents
Anthropic’s threat report shows AI moving from cyber assistant to operational loop
Anthropic’s September 2026 report shows Claude used in multi-step malicious workflows, shifting security attention from model outputs to operational loops.
-
Agents
GitHub turns Code Quality backlog fixes into a bulk Copilot agent job
GitHub Code Quality can assign up to 25 findings to Copilot at once, shifting remediation toward batch agent work while preserving pull-request review.
-
Industry
Red Hat AI 3.5 turns AI governance into platform infrastructure, with a support-tier catch
Red Hat AI 3.5 makes EvalHub GA and expands identity, GPU and agent controls, while some observability remains Technology Preview.