New Features & Tech Innovations
-
Open Source
FreeToken runs frontier-scale MoE models on consumer GPUs
FreeToken uses bandwidth-adaptive CPU-GPU execution and semantic caching to serve 35B to 753B MoE models on consumer and workstation hardware.
-
New Models
Qwen3.8-Max raises the bar for open-weight coding models
Alibaba's Qwen team says Qwen3.8-Max sets a new bar for coding and cowork tasks and marks its first open-weight Qwen-Max-class release.
-
Infrastructure & Chips
NVIDIA puts Groq 3 LPX into production for faster agent inference
NVIDIA has moved Groq 3 LPX into full production with Vera Rubin, targeting decode latency in token-heavy agent workloads and naming Nebius as the first cloud adopter.
-
Research
Inherent trains 27B Faraday agent to replicate research
Inherent’s 27B Faraday agent uses Replica and coding agents to reproduce scientific results, while benchmark caveats complicate direct model comparisons.
-
Infrastructure & Chips
Waymo reveals custom 5nm AI chip behind its robotaxi compute stack
Waymo reveals a custom 5nm ASIC for robotaxi sensor processing, showing a hybrid edge AI architecture built around specialized and merchant silicon.
-
New Models
Ant Ling opens six Ling 3.0 checkpoints for custom LLM training
Ant Ling releases six Ling 3.0 base checkpoints across two model sizes and three training stages for continued pretraining and specialized LLM work.