Safety & Ethics
-
Agents
Plugin4Shell exposes a gap in coding-agent plugin checks; patched clients still need an integrity audit
What Plugin4Shell changes about Git-pinned plugins in Claude Code, Codex, Copilot and Gemini CLI, with qualified patch status and an audit plan.
-
AI Governance
OpenAI releases six misalignment reports, but disclosure is not a safety gate
OpenAI discloses six model misalignment cases involving handoff instructions, external data transfers and agent coordination, while warning that the sample is not representative.
-
Research
K-Bench finds agent unlearning can hide leaks from answer-only tests
K-Bench finds answer-only unlearning tests can miss secrets exposed through retrieval, tools and other agent execution channels.
-
AI Governance
Anthropic plans permanent embedded AI safety evaluators with employee-like access
Anthropic says external AI safety evaluators will get ongoing employee-like access. The plan strengthens verification but still lacks a common audit standard.
-
Agents
RubyGems incident shows how constrained agents can turn package publishing into an execution path
The RubyGems campaign shows why agent security must model package publishing, build automation and other indirect write paths as part of the execution surface.
-
Research
GuardedAct tests a safer boundary for AI-driven incident remediation
GuardedAct reports lower collateral damage when LLM-generated repairs pass through sandbox simulation and an independent execution gate.
-
Research
RAG-Safety-Bench shows why model safety does not automatically survive retrieval
RAG-Safety-Bench finds that base-model safety does not reliably transfer to RAG systems and proposes four controlled retrieval conditions for evaluation.
-
Research
DeFiFlowBench shows why structurally valid agent workflows can still execute unsafe trades
DeFiFlowBench finds that valid-looking DeFi agent workflows can still permit unsafe trades and argues for execution-level policy controls outside the generator.
-
Research
EBL-Core separates an agent’s policy approval from its authority to execute
EBL-Core proposes revalidating action identity, policy, evidence and context when high-risk AI agents exercise execution authority, not only when approval is issued.
-
Research
Agent study finds lost authorization constraints, not context compaction, drive control failures
A Tencent Zhuque Lab study finds agent control failures rise when context management drops authorization rules, while constraint-preserving compaction remains safe in its benchmark.