INNER CODE UNIT · Python
create_multi_lifespan
cubist38/mlx-openai-server · app/server.py:358
def create_multi_lifespan(config: MultiModelServerConfig):
"""Create a FastAPI lifespan for multi-handler mode.
Each model entry in ``config.models`` is spawned in a dedicated
subprocess using ``multiprocessing.get_context("spawn")``, preventing
MLX Metal/GPU semaphore leaks (see
`<https://github.com/ml-explore/mlx/issues/2457>`_). A
``HandlerProcessProxy`` in the main process forwards requests to
the child via multiprocessing queues.
The proxies are registered in a ``ModelRegistry`` and attached to
``app.state.registry``. For backward compatibility the first
proxy is also stored as ``app.state.handler``.
Parameters
----------
config : MultiModelServerConfig
Parsed multi-model configuration (typically from YAML).