Aleph Alpha
Research

Aleph Alpha Blog

Introducing TFree-HAT 7B: Tokenizer-Free Models Achieving Top-Tier Multilingual Performance

Why Tokenizer-Free?

Benchmarks

LLM-as-a-Judge (MTBench)

Get Started

Hugging Face

pip install 'hat-splitter>=0.1.9' 'transformers==4.46.3' torch
pip install flash_attn
import torch
from transformers import AutoModelForCausalLM

INPUT = "When was Rome founded?"
MODEL_ID = "Aleph-Alpha/Llama-TFree-HAT-Pretrained-7B-DPO"

model = AutoModelForCausalLM.from_pretrained(
    trust_remote_code=True,
    pretrained_model_name_or_path=MODEL_ID,
    attn_implementation="flash_attention_2",
).to("cuda", torch.bfloat16)

input_ids, cumulative_word_lengths = model._prepare_input(INPUT, add_llama_template=True)
model_output = model.generate(
    input_ids,
    cumulative_seq_lengths_per_word=cumulative_word_lengths,
    max_new_tokens=300,
    use_cache=False,
)
print("Prompt: ", INPUT)
print("Completion: ", model_output.completion_text)

vLLM Fork

git clone https://github.com/Aleph-Alpha/vllm vllm-hat
cd vllm-hat

# Create and activate a 3.12 virtual env
uv venv -p 3.12
source .venv/bin/activate

# Tell vLLM to skip local compilation and use prebuilt CUDA wheels
export VLLM_USE_PRECOMPILED="1"

# Finally, install in editable mode
uv pip install -e

More blog posts