Safety & Ethics
-
Research
Lifecycle-hook updates can create an execution path outside agent guardrails
HookPry research shows why executable lifecycle-hook updates need permission-style review, least privilege and runtime provenance in agent platforms.
-
Research
Shared API drift can invalidate LLM-as-judge release gates
A preregistered audit finds shared LLM endpoints can undermine hard judge gates, pointing to an instrument-qualification step for production evaluation.
-
Research
PatchBench shows why a stopped crash is not a verified AI security fix
PatchBench finds that single-PoC checks can overstate AI vulnerability-patching success and argues for independent security and semantic validation.
-
Research
Context privilege escalation turns agent memory into a security boundary
A study across 12 agent harnesses shows how context can gain authority across roles and scopes, making provenance-aware non-escalation controls a runtime requirement.
-
New Models
GPT-6 Astra launches with critical cyber capability and staged access
OpenAI launches GPT-6 Astra with staged access, Critical cyber classification and runtime safeguards that can interrupt agent tasks.
-
New Models
OpenAI's Astra uses recurrent depth, raising new monitoring questions
Astra reportedly uses recurrent depth while OpenAI adds chain-of-thought monitoring. The design highlights why frontier AI needs independent runtime controls.
-
Agents
GitHub lets Copilot approvals count toward merge requirements
Copilot code review can now count as a required pull-request approval, with enterprise, repository and path controls around where AI approval applies.
-
Agents
GitHub extends Copilot content exclusions to app and CLI, but gaps remain
Copilot app and CLI now respect content exclusions, while GitHub still documents gaps in IDE agent modes, symlinks and remote filesystems.
-
Research
Cheap verifiers can make AI cascade dashboards blind to real errors
New research shows why LLM cascade routing and quality monitoring should not rely on the same verifier without an independent audit path.
-
Agents
CrowdStrike moves AI agent security from prompts to runtime execution
CrowdStrike Falcon Guardian links prompts and tool calls to endpoint actions, adding runtime controls for enterprise AI agents.