GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท gpt-oss-120b

GPT-OSS-120B

v20250805

OpenAI

Modelopen-sourceapache-2.0self-hostedprivacy
92
Exceptional
About This Model

OpenAI's first open-weight model released August 2025. 117B total params (5.1B active), Apache 2.0 license. Matches o4-mini on many benchmarks. Runs in 80GB memory.

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+

Flagship open-source performance. MoE architecture activates 5.1B of 117B params per token. Matches or beats o4-mini on most benchmarks.

task accuracy code

Competition coding and tool use benchmarks

Evidence
Codeforces Benchmark โ€” Matches o4-mini on competition coding
TauBench Tool Calling โ€” Exceeds o4-mini on tool calling
highVerified: 2026-07-09
task accuracy reasoning

Math competition benchmarks

Evidence
AIME 2024 & 2025 โ€” Outperforms o3-mini on competition mathematics
Chain-of-Thought Access โ€” Full chain-of-thought reasoning process exposed
highVerified: 2026-07-09
task accuracy general

General knowledge and domain-specific testing

Evidence
MMLU & HLE โ€” Matches o4-mini on general problem solving
HealthBench โ€” Exceeds o4-mini on health-related queries
highVerified: 2026-07-09
output consistency

Internal testing

Evidence
OpenAI Model Card โ€” Configurable reasoning effort for consistency
mediumVerified: 2026-07-09
latency p50

Median latency estimation

Evidence
Optimized Inference โ€” Fast inference with MoE architecture
mediumVerified: 2026-07-09
latency p95

95th percentile from community benchmarks

Evidence
Community Deployments โ€” ~2s on H100 hardware
mediumVerified: 2026-07-09
context window

Official specification

Evidence
OpenAI Technical Specs โ€” 128K context window natively supported
OpenAI Model Docs โ€” 131,072-token context window and up to 131,072 max output tokens
highVerified: 2026-07-09
uptime

Self-hosting provides full control

Evidence
Self-Hosted Model โ€” 100% uptime when self-hosted
highVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+

Good base security. Self-hosting provides complete control over safety guardrails and data handling. Note: Feb 2026 research (arXiv 2602.14689) shows open-weight models broadly remain vulnerable to prefill attacks โ€” pair self-hosted deployments with external guardrails. OpenAI's gpt-oss-safeguard (Oct 2025), an Apache-2.0 safety-classifier fine-tune of gpt-oss, can serve as a policy-based moderation layer.

prompt injection resistance

OWASP LLM01 testing

Evidence
OpenAI Safety Testing โ€” Good resistance, customizable for self-hosted
mediumVerified: 2026-07-09
jailbreak resistance

Adversarial testing

Evidence
Community Testing โ€” Standard resistance, self-host allows custom guardrails
Prefill Attack Study (arXiv 2602.14689) โ€” Large empirical study finds prefill attacks consistently effective against all major contemporary open-weight models; large reasoning models show partial resistance but remain vulnerable to tailored strategies
mediumVerified: 2026-07-09
data leakage prevention

Self-hosting analysis

Evidence
Self-Hosted Deployment โ€” Complete data control when self-hosted
highVerified: 2026-07-09
output safety

Safety testing

Evidence
OpenAI Safety โ€” Standard safety training, customizable
mediumVerified: 2026-07-09
api security

Deployment security review

Evidence
Self-Hosted Security โ€” Customer controls all API security when self-hosted
highVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+

Perfect privacy when self-hosted. No data sent to OpenAI. Full compliance control. Ideal for regulated industries.

data residency

Self-hosting analysis

Evidence
Open-Weight Model โ€” Deploy anywhere, full data residency control
highVerified: 2026-07-09
training data optout

Privacy model analysis

Evidence
Self-Hosted Model โ€” No data sent to OpenAI when self-hosted
highVerified: 2026-07-09
data retention

Self-hosting review

Evidence
Self-Hosted Deployment โ€” Complete control over data retention
highVerified: 2026-07-09
pii handling

Data flow analysis

Evidence
On-Premises Deployment โ€” PII never leaves your infrastructure
highVerified: 2026-07-09
compliance certifications

Compliance model review

Evidence
Self-Hosted Compliance โ€” Inherit your infrastructure's certifications
highVerified: 2026-07-09
zero data retention

Privacy architecture review

Evidence
Open-Weight Model โ€” Complete control, zero external retention
highVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+

Exceptional transparency. Full chain-of-thought access. Complete model weights and architecture disclosed. Open-source enables auditing.

explainability

Reasoning transparency

Evidence
Full Chain-of-Thought โ€” Complete access to reasoning process
highVerified: 2026-07-09
hallucination rate

QA testing

Evidence
Benchmark Testing โ€” Good factual accuracy
mediumVerified: 2026-07-09
bias fairness

Bias benchmarks

Evidence
OpenAI Model Card โ€” Standard bias testing
mediumVerified: 2026-07-09
uncertainty quantification

Confidence assessment

Evidence
Model Behavior โ€” Good uncertainty expression
mediumVerified: 2026-07-09
model card quality

Documentation review

Evidence
Comprehensive Model Card โ€” Detailed technical specs, benchmarks, architecture
highVerified: 2026-07-09
training data transparency

Training data disclosure review

Evidence
OpenAI Documentation โ€” Mostly English, STEM, coding focus disclosed
highVerified: 2026-07-09
guardrails

Safety mechanism review

Evidence
Customizable Guardrails โ€” Standard safety, customizable when self-hosted
highVerified: 2026-07-09
โš™๏ธOperational Excellence
+

Exceptional operational flexibility. Apache 2.0 enables commercial use. Massive deployment ecosystem. Self-host or use managed platforms.

api design quality

API compatibility review

Evidence
Deployment Platforms โ€” Works with vLLM, Ollama, llama.cpp, Azure, AWS, etc.
highVerified: 2026-07-09
sdk quality

SDK ecosystem review

Evidence
GitHub Repository โ€” Official repo, Hugging Face integration
highVerified: 2026-07-09
versioning policy

Version stability analysis

Evidence
Open Weights โ€” Weights frozen, no deprecation risk
Hugging Face Repository โ€” Repo last updated 2025-08-26; weights stable with no new revisions, remains freely downloadable under Apache 2.0
highVerified: 2026-07-09
monitoring observability

Monitoring capability review

Evidence
Self-Hosted Control โ€” Full observability when self-hosted
highVerified: 2026-07-09
support quality

Support ecosystem assessment

Evidence
Community Support โ€” GitHub issues, community forums, deployment partners
mediumVerified: 2026-07-09
ecosystem maturity

Ecosystem breadth analysis

Evidence
Deployment Partners โ€” Azure, Hugging Face, AWS, Fireworks, Together AI, Databricks, Vercel, Cloudflare, OpenRouter
highVerified: 2026-07-09
license terms

License review

Evidence
Apache 2.0 License โ€” Permissive Apache 2.0, no copyleft, no patent risk
highVerified: 2026-07-09
Strengths
  • +Apache 2.0 open-weight license enables commercial use without restrictions
  • +Matches or beats o4-mini on coding, math, and health benchmarks
  • +Complete data privacy when self-hosted (zero external data transmission)
  • +Full chain-of-thought reasoning access for transparency and debugging
  • +MoE architecture: 5.1B active of 117B total params, runs in 80GB
  • +Massive deployment ecosystem (Azure, AWS, Hugging Face, vLLM, Ollama)
Limitations
  • !Requires 80GB GPU memory (H100 or equivalent)
  • !Self-hosting complexity and infrastructure costs
  • !Community support vs enterprise SLA
  • !Slightly lower performance than flagship closed models
  • !No built-in safety guardrails (customizable but requires setup)
Metadata
pricing
input: Free (self-hosted)
output: Free (self-hosted)
notes: Infrastructure costs only: ~$2-4/hr for H100. Managed platforms vary. Free for download and commercial use under Apache 2.0.
last verified: 2026-07-09
context window: 131072
languages
0: English
1: Spanish
2: French
3: German
4: Italian
5: Portuguese
6: Japanese
7: Korean
8: Chinese
modalities
0: text
api endpoint: Self-hosted (various platforms)
model download: https://huggingface.co/openai/gpt-oss-120b
github: https://github.com/openai/gpt-oss
open source: true
license: Apache 2.0
architecture: Mixture-of-Experts (MoE) Transformer
parameters: 117B total (5.1B active per token)
memory requirement: 80GB (MXFP4 quantization)
tokenizer: o200k_harmony
deployment platforms
0: Azure
1: AWS
2: Hugging Face
3: vLLM
4: Ollama
5: llama.cpp
6: LM Studio
7: Fireworks
8: Together AI
9: Baseten
10: Databricks
11: Vercel
12: Cloudflare
13: OpenRouter

Use Case Ratings

code generation

Excellent coding. Matches o4-mini. Configurable reasoning effort. Full chain-of-thought debugging.

customer support

Good for customer support. Self-host for complete data privacy. Configurable reasoning for cost control.

content creation

Strong content creation. Self-hosting enables unlimited generation without API costs.

data analysis

Excellent for data analysis. Keep sensitive data on-premises. Full chain-of-thought for transparency.

research assistant

Outstanding for research. 128K context. Self-host proprietary research data. Full reasoning transparency.

legal compliance

Perfect for legal. Self-host for complete compliance. No data leaves premises. Apache 2.0 license clarity.

healthcare

Ideal for healthcare. Self-host for HIPAA. Complete PHI privacy. No external data transmission.

financial analysis

Excellent for finance. Outperforms o3-mini on math. Self-host proprietary financial data.

education

Great for education. Full chain-of-thought shows reasoning steps. Self-host for institutional control.

creative writing

Good creative writing. Unlimited generation when self-hosted. No API costs for iteration.