INNER CODE UNIT · Python
load_gguf_hf_model
junruxiong/IncarnaMind · toolkit/local_llm.py:23
def load_gguf_hf_model(
model_id: str,
model_basename: str,
max_tokens: int,
temperature: float,
device_type: str,
):
"""
Load a GGUF/GGML quantized model using LlamaCpp.
This function attempts to load a GGUF/GGML quantized model using the LlamaCpp library.
If the model is of type GGML, and newer version of LLAMA-CPP is used which does not support GGML,
it logs a message indicating that LLAMA-CPP has dropped support for GGML.
Parameters:
- model_id (str): The identifier for the model on HuggingFace Hub.
- model_basename (str): The base name of the model file.
- max_tokens (int): The maximum number of tokens to generate in the completion.