FLEET turns repeated LLM sampling into memory-guided search
FLEET reports better LiveCodeBench Pass@32 and roughly 3× sample-efficiency by adding memory to repeated LLM decoding, with important deployment limits.
FLEET reports better LiveCodeBench Pass@32 and roughly 3× sample-efficiency by adding memory to repeated LLM decoding, with important deployment limits.
Microsoft adds Copilot Home, Code, Autopilot and Managed Runtime, creating a governed execution layer for persistent workplace agents.
Foundry Routines now provide managed scheduled and event-driven agent execution with identity selection, run history and continuation patterns.
Rollouts and Security Review extend Cursor agents from pull requests into deployment health and exploit-path security review.
AIDE² repeatedly rewrote its own research-agent harness in an eight-day experiment. Held-out tests suggest transfer, while the evaluation gate remains the key control boundary.
Anthropic says Claude agents found a previously uncharacterized repeat-associated enzyme system. Wet-lab evidence is promising, but ART's biological function remains unknown.
DeepSeek Harness 0.1.7 makes agent capabilities more dynamically composable. Aipolix examines the benefits and the new operational risks.
GPT-6 Sol and Luna introduce a 20x price spread, large context and different cost boundaries. Aipolix examines agent routing economics.
Claude Opus 5.5 cuts token and cache costs while sensitive requests may use fallback models. Aipolix examines the impact on agent evaluation.
Grok 4.7 keeps basic API token rates unchanged, but independent tests show higher output consumption and provider-specific long-context pricing.