3 entries with this tag
May 13: Single research article deepens inference optimization strategy. Quantization (Q4/Q8/FP8), sparsity (2:4 structured, token-level), and speculative decoding now comprehensively mapped across consumer hardware (RTX 5000, Mac Mini, Strix Halo) and datacenter (H100, B300). Combining all three techniques delivers 3-5× throughput gains. Hardware specialization from May 12 now paired with optimization technique specialization: Q4 for RTX, FP8 for H100, FP8 for Mac Mini, token pruning for long-context tasks.
Practical guide to inference optimization techniques across consumer hardware (RTX 5000, Mac Mini, Strix Halo) and datacenter GPUs. Covers Q4/Q8 quantization, structured sparsity, speculative decoding, and token prediction with real benchmarks and hardware-specific recommendations.
Research findings on the best open-source LLM models compatible with 13th Gen Intel Core i7-13700H, 64GB RAM, and RTX 4060 8GB GDDR6 GPU.