ApX logoApX logo

Command A

Parameters

111B

Context Length

256K

Modality

Text

Architecture

Dense

License

CC-BY-NC

Release Date

13 Mar 2025

Knowledge Cutoff

Jun 2024

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

234.74 GB VRAM

Consumer

12x RTX 4090

24GB VRAM

Datacenter

4x NVIDIA A100

80GB VRAM

Apple Silicon

3x Apple M3 Max

128GB VRAM

256,000 tokens

269.83 GB VRAM

Consumer

14x RTX 4090

24GB VRAM

Datacenter

4x NVIDIA A100

80GB VRAM

Apple Silicon

3x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteContext: 256Kx N layersNormPre-AttentionMulti-Head Attention+NormPre-FFNFeed-Forward NetworkSwiGLU+Final NormOutput Logits

Evaluation Benchmarks

Rank

#152

BenchmarkScoreRank

0.858

10

0.12

26

0.229

30

Professional Knowledge

MMLU Pro

0.69

40

General Text

Text Arena

1354

125

Intelligence Index

Artificial Analysis

0.07

127

Rankings

Overall Rank

#152

Coding Rank

#110

About Command A

Cohere Command A is a large-scale generative model engineered for high-performance enterprise workflows, particularly those involving tool use, agents, and retrieval-augmented generation (RAG). Developed to provide a high-throughput alternative for production environments, the model maintains a significant parameter count of 111 billion while optimizing for deployment on common dual-GPU hardware configurations. Its design focuses on business-critical accuracy and speed, supporting a standard context window of 256,000 tokens to facilitate the processing of extensive corporate documentation and long-form conversational histories.

The model's architecture is a decoder-only transformer that utilizes several sophisticated structural innovations to balance local and global context. It features a hybrid attention mechanism where three-quarters of the layers employ sliding window attention for efficient local modeling, while every fourth layer uses a full global attention mechanism to maintain long-range dependencies. Technical specifications include the use of Grouped Query Attention (GQA) for optimized inference throughput, SwiGLU activation functions for improved gradient flow, and the omission of bias terms to stabilize the training process. Positional information is handled via Rotary Positional Embeddings (RoPE) in local attention layers, whereas global layers utilize a position-agnostic approach.

Optimized for global enterprise deployment, Command A is trained across 23 languages, including major business languages such as English, French, Spanish, Chinese, and Arabic. The model is specifically aligned for conversational tool use, allowing it to interact with external APIs, databases, and search engines with high precision. This alignment, achieved through supervised fine-tuning and preference optimization, makes it particularly effective for multi-step agentic reasoning and financial data manipulation where extracting numerical details from complex contexts is required.

Technical Specifications

Attention

Attention Structure

Multi-Head Attention

Attention Heads

-

Key-Value Heads

-

Attention Head Dimension

-

Position Embedding

Absolute Position Embedding

RoPE Theta

-

Sliding Window Attention

-

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

-

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

-

Number of Layers

-

FFN Intermediate Size (Dense)

-

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

-

About Command


Other Command Models
Command A: Specifications and GPU VRAM Requirements