ApX logoApX logo

Optimus Alpha

Parameters

Undisclosed

Context Length

128K

Modality

Multimodal

Architecture

Undisclosed

License

Proprietary

Release Date

10 Jan 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Input: $2.00 · Output: $8.00

Evaluation Benchmarks

BenchmarkScoreRank

Coding

Archived
Aider Coding

0.53

15

Rankings

Overall Rank

-

Coding Rank

-

About Optimus Alpha

NVIDIA Optimus Alpha delivers optimized AI inference with a focus on efficiency and throughput. Features hardware-aware optimizations for NVIDIA GPUs, enabling high-performance deployment in enterprise environments. Excels at sustained high-throughput workloads with consistent low latency. Ideal for production deployments requiring reliable performance at scale on NVIDIA infrastructure.

Technical Specifications

Architecture specifications are undisclosed for proprietary models.

Attention

Attention Structure

Multi-Head Attention

Attention Heads

Undisclosed

Key-Value Heads

Undisclosed

Attention Head Dimension

Undisclosed

Position Embedding

Absolute Position Embedding

RoPE Theta

Undisclosed

Sliding Window Attention

Undisclosed

Sliding Window Size

Undisclosed

Sliding Window Ratio

Undisclosed

Linear Attention

Undisclosed

Linear Attention Ratio

Undisclosed

Normalization

Undisclosed

Activation Function

Undisclosed

Dimensions

Hidden Dimension Size

Undisclosed

Number of Layers

Undisclosed

FFN Intermediate Size (Dense)

Undisclosed

Multi-Token Prediction Heads

Undisclosed

Tokenizer

Vocabulary Size

Undisclosed

About Optimus

NVIDIA's Optimus Alpha models combine advanced AI capabilities with hardware-software co-optimization. Built for enterprise deployments requiring high throughput, low latency, and efficient resource utilization on NVIDIA infrastructure.


Other Optimus Models
  • No related models available
Optimus Alpha: Model Specifications and Details