Parameters
1.3B
Context Length
2K
Modality
Text
Architecture
Dense
License
Apache-2.0
Release Date
29 Feb 2024
Knowledge Cutoff
Nov 2023
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
1x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
2,048 tokens
Consumer
1x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
No evaluation benchmarks for CroissantLLM Base available.
Overall Rank
-
Coding Rank
-
CroissantLLM Base is a 1.3 billion parameter decoder-only transformer model designed to provide balanced bilingual proficiency in French and English. Unlike many contemporary large language models that treat non-English languages as secondary through minor data inclusion, CroissantLLM was pre-trained using a strictly balanced 1:1 ratio of French and English data. This architectural choice aims to mitigate linguistic bias and ensure that French cultural and technical knowledge is represented with the same fidelity as English. The model was trained on 3 trillion tokens, a substantial corpus that exceeds the training volume of many larger open-source models in its class.
Technically, the model is built upon the Llama architecture, incorporating established components such as Rotary Positional Encodings (RoPE) and RMSNorm to stabilize deep network activations. To optimize for the bilingual use case, the developers introduced a custom SentencePiece-based tokenizer trained on a high-quality mix of French, English, and code data. This tokenizer achieves significantly lower fertility rates for French text compared to standard multilingual tokenizers, improving both computational efficiency and the model's ability to capture linguistic nuances. The architecture features 24 layers with a hidden dimension of 2048 and 16 attention heads, following a dense structure without the use of mixture-of-experts.
CroissantLLM Base is engineered for high performance on consumer-grade hardware, making it suitable for deployment on local devices such as personal computers and mobile systems. Its training history is highly transparent, with the researchers releasing extensive details on the pre-training data and providing access to checkpoints throughout the training process. The model serves as a foundation for various downstream tasks, particularly translation and content generation in French-centric environments, where its specialized vocabulary and balanced training provide a distinct advantage over models trained on predominantly English-centric datasets.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
16
Key-Value Heads
16
Attention Head Dimension
128
Position Embedding
Absolute Position Embedding
RoPE Theta
10,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
2,048
Number of Layers
24
FFN Intermediate Size (Dense)
5,504
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
32,000
CroissantLLM is a bilingual French-English language model developed by French research institutions. The model is trained on a curated mix of French and English data to provide language understanding while preserving French linguistic heritage. It is designed for low-resource inference on consumer-grade hardware.
APX AI
Online