INNER CODE UNIT · Python

load_gguf_hf_model

junruxiong/IncarnaMind · toolkit/local_llm.py:23

def load_gguf_hf_model(
    model_id: str,
    model_basename: str,
    max_tokens: int,
    temperature: float,
    device_type: str,
):
    """
    Load a GGUF/GGML quantized model using LlamaCpp.

    This function attempts to load a GGUF/GGML quantized model using the LlamaCpp library.
    If the model is of type GGML, and newer version of LLAMA-CPP is used which does not support GGML,
    it logs a message indicating that LLAMA-CPP has dropped support for GGML.

    Parameters:
    - model_id (str): The identifier for the model on HuggingFace Hub.
    - model_basename (str): The base name of the model file.
    - max_tokens (int): The maximum number of tokens to generate in the completion.

View source record →

📰 Research Paper
Loading…
⏳ Fetching content…