J
My single-GPU vLLM serving setup — Dockerfile and config included
Been running a 70B Q4 on one 80GB card for a side project. Key tricks: prefix caching on, max_num_seqs tuned, and a tiny health check for the load balancer. Sharing my config because I wish someone had shared theirs.
댓글 0
첫 댓글을 남겨보세요.
로그인 후 댓글을 쓸 수 있어요.