Open-source C++/CUDA inference engine for ternary-weight LLMs with runtime LoRA. Solves the ternary merge problem: LoRA deltas (~1e-5 magnitude) are silently erased when merged into ternary weights...
Pantheon runs about fifty targeted workloads against NVIDIA and AMD GPUs to find out whether a card is healthy: memory diagnostics adapted from the DRAM testing literature (march tests, disturb pat...