Research
K-Bench finds agent unlearning can hide leaks from answer-only tests
K-Bench finds answer-only unlearning tests can miss secrets exposed through retrieval, tools and other agent execution channels.
K-Bench finds answer-only unlearning tests can miss secrets exposed through retrieval, tools and other agent execution channels.
TAM tests long professional procedures. GPT-5 baselines reached only 1% exact match on ICD coding and 15.5% on sentencing.