INNER CODE UNIT · Python

create_multi_lifespan

cubist38/mlx-openai-server · app/server.py:358

def create_multi_lifespan(config: MultiModelServerConfig):
    """Create a FastAPI lifespan for multi-handler mode.

    Each model entry in ``config.models`` is spawned in a dedicated
    subprocess using ``multiprocessing.get_context("spawn")``, preventing
    MLX Metal/GPU semaphore leaks (see
    `<https://github.com/ml-explore/mlx/issues/2457>`_).  A
    ``HandlerProcessProxy`` in the main process forwards requests to
    the child via multiprocessing queues.

    The proxies are registered in a ``ModelRegistry`` and attached to
    ``app.state.registry``.  For backward compatibility the first
    proxy is also stored as ``app.state.handler``.

    Parameters
    ----------
    config : MultiModelServerConfig
        Parsed multi-model configuration (typically from YAML).

View source record →

📰 Research Paper
Loading…
⏳ Fetching content…