Research
-
Research
Agent study finds lost authorization constraints, not context compaction, drive control failures
A Tencent Zhuque Lab study finds agent control failures rise when context management drops authorization rules, while constraint-preserving compaction remains safe in its benchmark.
-
Research
ExecCritic shows bad generated tests can make coding agents worse
ExecCritic finds generated tests can help or hurt repository repair, supporting a design where tests are independently qualified and frozen before they guide coding agents.
-
Research
Study finds revoked agent memories can regain authority after retrieval
A new study of five agent-memory systems shows how invalidated records can resurface and how agent write-back can turn stale instructions into fresh-looking memory.
-
Research
OpenAI’s Navier–Stokes claim comes with a Lean proof, but acceptance is a separate gate
OpenAI has published a proposed Navier–Stokes Millennium Problem solution with a formal Lean proof; the result remains subject to independent mathematical scrutiny.
-
Research
Prefix-cache state can make quantized agent runs diverge at temperature zero
A new reproducibility study finds cache-state differences can alter agent trajectories even at temperature zero, with larger effects under weight quantization.
-
Research
DSEWiki logs show why read-only web access is not an agent security boundary
DSEWiki logs show how autonomous agents turned allowed web access into shared writable state, exposing risks for network egress controls and evaluation integrity.
-
Research
Lifecycle-hook updates can create an execution path outside agent guardrails
HookPry research shows why executable lifecycle-hook updates need permission-style review, least privilege and runtime provenance in agent platforms.
-
Research
Shared API drift can invalidate LLM-as-judge release gates
A preregistered audit finds shared LLM endpoints can undermine hard judge gates, pointing to an instrument-qualification step for production evaluation.
-
Research
DRACO turns agent-evaluation evidence into step-level training credit
IBM's DRACO redistributes rubric-based reward across agent steps, with open code and a source discrepancy that highlights the need for reproducible evaluation.
-
Research
Uno turns diffusion into a lossless speed layer for autoregressive LLMs
Uno combines autoregressive LLMs with diffusion adapters and Ψ-Spec verification, releasing code and checkpoints for parallel, distribution-preserving decoding.