Aleph Alpha
Research

Aleph Alpha Blog

Introducing Pharia-1-LLM: transparent and compliant

Pharia-1-LLM-7B

Dataset

Model architecture & hyperparameters

Component GPT Llama 2
BiasTrueFalse
NormLayernormRMS Norm
MLPStandard FFNSwiGLU
Pharia-1-LLM-7B-control, Pharia-1-LLM-7B-control-aligned
Number of layers27
Number of attention heads36
Head size128
Number of Key-Value heads4
Hidden size4608
MLP expansion factor4
MLP typeStandard
Vocabulary size128,000
Rotary base1,000,000
Dropout0
Weight decay0.1
Learning rate scheduleWarmup: Linear over 2000 steps · Max value: 3.0e-4 · Decay: Cosine to 0.0
Total parameter count7,041,544,704

Pre-training

Model GPUs Pipeline-parallel size Model-parallel size Data-parallel size Batch size Micro-batch size
Pharia-1-LLM-7B2562112810241
Model name Avg. step duration Avg. MFU
Pharia-1-LLM-7B8.6s (A100) · 3.6s (H100)0.66 (A100) · 0.5 (H100)

Fine-tuning

Pharia-1-LLM-7B-control
Learning rate scheduleWarmup: Linear over 100 steps · Max value: 1.0e-6 · Decay: Cosine to 5e-09
Batch size256
Train steps (number of epochs)12500 (~ 5)
Dataset size~ 640,000
Pharia-1-LLM-7B-control
Learning rate scheduleWarmup: Linear over 100 steps · Max value: 5.0e-7 · Decay: Cosine to 5e-09
Batch size8
Train steps (number of epochs)1200 (3)
Dataset size3,200

Evaluation

LLM Use-case Example Evaluation Task Example Evaluation Task Source Real Use-Case Equivalent Task Explanation
Domain-specific logical reasoningWhich statement best explains the purpose of Hart's distinction between 'being obliged' and 'having an obligation'? Output is expected to be one of: 'It demonstrates the difference between the internal and the external aspect of a rule.', 'It refutes the natural lawyer view of the role of morality in law', 'It explains the nature of power-conferring rules.', 'It illuminates the concept of a rule.'MMLUYou are given a legal text and a contract closed between (party_a) and (party_b). Reason step by step to determine whether the contract actually fulfills the requirements stated in the legal text. Legal text: (legal_text); Contract: (contract)Domain specific reasoning tasks often lean on world knowledge acquired by the model during training. In practice, however, application will always want to inject one or more contexts to ground the reasoning and reduce the risk of hallucination. However, such evaluation examples are difficult to come by. They often lean on proprietary data (e.g. contracts) safeguarded by industry. Even if such data is available, deep domain knowledge is still necessary to construct proper, useful instruction scenarios. This is usually a main cost factor in data creation & acquisition. Evaluation is resource-intensive as the output can only be verified by domain experts. Judging with an LLM is not sufficient in such cases.
SummarizationSummarize the article you have been given in a brief manner. Mathematics and art are related in a variety of ways. [...]Alpaca Eval(document); Given the above document, write a summary which has no more than 3 sentences long that also satisfies the following constraints: (list_of_constraints)Evaluation frameworks geared towards assessing general instructability tend to ignore the need for a variety of requirements and constraints that users want to impose onto the completion. The main reasons are: Difficulty in learning about such cases as such requirements may be safeguarded by the organizations dictating them. High costs associated with constructing such samples. Complexity in deciding an evaluation criterion that measures the achieved success granularly. Extra complexity and cost associated with properly evaluating whether all requirements/constraints were actually fulfilled.

More blog posts