Agents
-
Research
Prefix-cache state can make quantized agent runs diverge at temperature zero
A new reproducibility study finds cache-state differences can alter agent trajectories even at temperature zero, with larger effects under weight quantization.
-
Research
DSEWiki logs show why read-only web access is not an agent security boundary
DSEWiki logs show how autonomous agents turned allowed web access into shared writable state, exposing risks for network egress controls and evaluation integrity.
-
Agents
OpenClaw 2026.9.2 makes cross-agent session access the default
OpenClaw now enables Gateway-wide session visibility and agent-to-agent access by default, making trust-boundary review part of multi-agent upgrades.
-
Research
Lifecycle-hook updates can create an execution path outside agent guardrails
HookPry research shows why executable lifecycle-hook updates need permission-style review, least privilege and runtime provenance in agent platforms.
-
New Products
NVIDIA PAIR scales local agent inference by routing requests, not pooling VRAM
NVIDIA PAIR can widen local agent throughput across PCs and Macs, but it routes independent requests rather than pooling VRAM or sharding models.
-
Agents
VS Code Agent Merge turns PR blockers into a continuous agent loop
VS Code 1.136 lets agents repeatedly fix PR reviews, CI failures and conflicts, raising a new need for independent merge controls.
-
Research
DRACO turns agent-evaluation evidence into step-level training credit
IBM's DRACO redistributes rubric-based reward across agent steps, with open code and a source discrepancy that highlights the need for reproducible evaluation.
-
Research
PatchBench shows why a stopped crash is not a verified AI security fix
PatchBench finds that single-PoC checks can overstate AI vulnerability-patching success and argues for independent security and semantic validation.
-
Research
Context privilege escalation turns agent memory into a security boundary
A study across 12 agent harnesses shows how context can gain authority across roles and scopes, making provenance-aware non-escalation controls a runtime requirement.
-
🧩 AI Agents Need Contracts, Not Better Prompts (5/10)
How versioned agent contracts define scope, tools, budgets, refusal, escalation and evidence, and turn AI governance into enforceable runtime behavior.