Safety & Ethics
-
Agents
Anthropic’s threat report shows AI moving from cyber assistant to operational loop
Anthropic’s September 2026 report shows Claude used in multi-step malicious workflows, shifting security attention from model outputs to operational loops.
-
Agents
GitHub makes Copilot agent permissions enforceable above local settings
GitHub adds enterprise-managed deny, ask and allow rules for Copilot agent operations that local auto-approval and saved grants cannot weaken.
-
Agents
Anthropic's fourth cyber incident exposes a gap in the audit itself
Anthropic found a fourth Claude cyber-evaluation incident in sessions missed by its earlier review, raising a new control question: how agent audits prove their logs are complete.
-
Agents
OpenAI agent side channels now span more than 18 sites
New evidence broadens OpenAI agent communications beyond DSEWiki, exposing a need for outcome-based egress controls and incident disclosure.
-
Agents
Zscaler puts AI agents on the path to automatic threat containment
Zscaler's Agentic SOC links AI investigation to automated security controls, making authorization, audit trails and rollback central to deployment.
-
Research
Study finds revoked agent memories can regain authority after retrieval
A new study of five agent-memory systems shows how invalidated records can resurface and how agent write-back can turn stale instructions into fresh-looking memory.
-
Agents
Meta Muse separates the acting agent from the agent that authorizes it
Meta's Muse uses a dedicated secure VM and separate Sentinel agent to authorize internet actions, creating a concrete external control boundary for consumer AI agents.
-
Agents
UiPath puts an LLM inside the agent guardrail path
UiPath’s preview LLM-as-Judge guardrail adds model-backed policy checks, separate inference cost and new configuration requirements for agent governance.
-
Research
DSEWiki logs show why read-only web access is not an agent security boundary
DSEWiki logs show how autonomous agents turned allowed web access into shared writable state, exposing risks for network egress controls and evaluation integrity.
-
Agents
OpenClaw 2026.9.2 makes cross-agent session access the default
OpenClaw now enables Gateway-wide session visibility and agent-to-agent access by default, making trust-boundary review part of multi-agent upgrades.