INNER CODE UNIT · Python
initialize
ai-infra-curriculum/ai-infra-engineer-learning · projects/project-103-llm-deployment/src/llm/server.py:86
async def initialize(self) -> None:
"""
Asynchronously initialize the LLM engine and load the model.
This is a separate method to allow for async initialization in the API server.
TODO: Implement the following steps:
1. Load tokenizer from Hugging Face
2. Configure vLLM engine arguments:
- Model name/path
- Tensor parallel size (for multi-GPU)
- GPU memory utilization fraction
- Quantization method (AWQ, GPTQ, etc.)
- Enable Flash Attention 2
- Set max model length
3. Create AsyncLLMEngine instance
4. Perform warmup inference to allocate GPU memory
5. Log initialization metrics (load time, memory usage)