Open Source
Colibrì runs DeepSeek V4.1 Flash from disk, but I/O is still the real limit
Colibrì 1.11.0 adds native DeepSeek V4.1 Flash support, showing how disk-tiered inference trades memory capacity for I/O and latency.
Colibrì 1.11.0 adds native DeepSeek V4.1 Flash support, showing how disk-tiered inference trades memory capacity for I/O and latency.
Cohere releases North Small Translate with downloadable weights, 50-language support and MoE serving options, but public weights remain non-commercial.
Tencent releases Hy4 preview with Apache 2.0 weights, 1M-token context, vLLM/SGLang deployment and a coding-focused 770B MoE architecture.
FreeToken uses bandwidth-adaptive CPU-GPU execution and semantic caching to serve 35B to 753B MoE models on consumer and workstation hardware.