
Multiverse Computing's AI model compression boosts Llama 3.3 70B performance on Intel Xeon 6 CPUs by nearly 94%.
Multiverse Computing announced that its CompactifAI-compressed Llama 3.3 70B AI model now runs efficiently on Intel Xeon 6 processors, nearly doubling throughput and halving latency compared to the uncompressed version. This breakthrough reduces the ...



