Open Source
FreeToken runs frontier-scale MoE models on consumer GPUs
FreeToken uses bandwidth-adaptive CPU-GPU execution and semantic caching to serve 35B to 753B MoE models on consumer and workstation hardware.
FreeToken uses bandwidth-adaptive CPU-GPU execution and semantic caching to serve 35B to 753B MoE models on consumer and workstation hardware.