Website: projectdmx.github.io
- DMI — A decoupled, asynchronous observation substrate for high-speed LLM inference.
Website: projectdmx.github.io
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
A high-throughput and memory-efficient inference and serving engine for LLMs
Ongoing research training transformer models at scale
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM
This organization has no public members. You must be a member to see who’s a part of this organization.
Loading…
Loading…