2 entries with this tag
June 30: One major research article — DeepSeek V4 and DSpark, the open-source efficiency breakthrough with 1.6T MoE, 1M context, and 85% faster inference via speculative decoding.
Practical guide to inference optimization techniques across consumer hardware (RTX 5000, Mac Mini, Strix Halo) and datacenter GPUs. Covers Q4/Q8 quantization, structured sparsity, speculative decoding, and token prediction with real benchmarks and hardware-specific recommendations.