GitHub / InseeFrLab / auto-tuning-vllm
Auto-tuning for vllm. Getting the best performance out of your LLM deployment (vllm+guidellm+optuna)
JSON API: https://ecosystem.code.gouv.fr/api/v1/hosts/GitHub/repositories/InseeFrLab%2Fauto-tuning-vllm
Stars: 6
Forks: 0
Open issues: 15
License: apache-2.0
Language: Python
Size: 2.99 MB
Dependencies parsed at: Pending
Created at: 6 months ago
Updated at: 9 days ago
Pushed at: 15 days ago
Last synced at: 5 days ago
Topics: benchmarking, gpu, guidellm, hyperparameter-optimization, llm-inference, llm-serving, llmops, optuna, speculative-decoding, vllm