LLMs1 min read
X-CoSD: Efficient Cross-Vocabulary LLM Inference
X-CoSD is a new distributed inference framework for large language models that reduces communication load by handling heterogeneous vocabularies. Experiments show improved token generation speed with comparable generation quality to the server LLM.
From arXiv cs.CL