Neural Magic
Neural Magic (Acquired by Red Hat) empowers developers to optimize & deploy LLMs at scale. Our model compression & acceleration enable top performance with vLLM
Pinned Loading
仓库
Showing 10 of 105 repositories
- vllm Public 复刻ed from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- nm-vllm-omni-ent Public 复刻ed from vllm-project/vllm-omni
A framework for efficient model inference with omni-modality models
- inspect_k8s_sandbox Public 复刻ed from UKGovernmentBEIS/inspect_k8s_sandbox
A Kubernetes sandbox environment for use with inspect_ai
- DeepGEMM Public 复刻ed from deepseek-ai/DeepGEMM
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
- vllm-envs Public
Fast disposable vLLM dev environments from any branch/commit, built on content-addressed layer caches
- DeepEP Public 复刻ed from deepseek-ai/DeepEP
DeepEP: an efficient expert-parallel communication library
Top languages
Loading…
Most used topics
Loading…