GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท openai-o1

OpenAI o1

v20250915

OpenAI

Modeldeprecatedreasoningchain-of-thoughtcoding
89
Strong
About This Model

DEPRECATED: o1 variants and o1-pro shut down in the API on 2026-10-23; already removed from ChatGPT (o1-preview/o1-mini removed in 2025). Migration target is GPT-5.5. Historically an advanced reasoning model (57.1% SWE-bench, 79.2% HumanEval) with extended chain-of-thought reasoning.

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+

Exceptional reasoning capabilities with extended chain-of-thought. Best for complex problem-solving requiring deep thinking. Higher latency due to reasoning overhead.

task accuracy code

Industry-standard coding benchmarks measuring real-world software engineering tasks

Evidence
SWE-bench Verified โ€” 57.1% resolution rate
HumanEval โ€” 79.2% accuracy on code generation
highVerified: 2026-07-09
task accuracy reasoning

Competition-level reasoning benchmarks requiring extended chain-of-thought

Evidence
AIME 2024 โ€” 83% on high school competition math (top 500 US students)
GPQA Diamond โ€” 78.3% on PhD-level science questions
highVerified: 2026-07-09
task accuracy general

Comprehensive knowledge testing across domains

Evidence
MMLU โ€” 85.5% on comprehensive knowledge benchmark
OpenAI Benchmarks โ€” Strong general performance with reasoning optimization
highVerified: 2026-07-09
output consistency

Internal testing with repeated prompts at various temperature settings

Evidence
OpenAI Documentation โ€” High consistency due to chain-of-thought reasoning
highVerified: 2026-07-09
latency p50

Median latency for API requests with standard prompt sizes

Evidence
OpenAI Documentation โ€” Typical response time ~4.5s due to extended reasoning
highVerified: 2026-07-09
latency p95

95th percentile response time across diverse workloads

Evidence
Community benchmarking โ€” p95 latency ~8.2s for complex reasoning
highVerified: 2026-07-09
context window

Official specification from provider

Evidence
OpenAI Documentation โ€” 128K token context window
highVerified: 2026-07-09
uptime

Historical uptime data from official status page

Evidence
OpenAI Status Page โ€” 99.9% uptime (last 90 days)
highVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+

Strong security posture with enhanced reasoning-based safety. Good protection against common attacks.

prompt injection resistance

Testing against OWASP LLM01 prompt injection attacks

Evidence
OpenAI Safety Research โ€” Strong resistance to prompt injection
highVerified: 2026-07-09
jailbreak resistance

Testing against adversarial prompt datasets

Evidence
OpenAI Safety Evaluations โ€” Enhanced jailbreak resistance via chain-of-thought
highVerified: 2026-07-09
data leakage prevention

Analysis of privacy policies and data handling practices

Evidence
OpenAI Privacy Policy โ€” No training on API data without opt-in
mediumVerified: 2026-07-09
output safety

Comprehensive safety testing across harmful content categories

Evidence
OpenAI Safety Evaluations โ€” Comprehensive safety filtering
highVerified: 2026-07-09
api security

Review of API security features and best practices

Evidence
OpenAI API Documentation โ€” API key authentication, OAuth, HTTPS, rate limiting
highVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+

Good privacy posture with SOC 2 certification. 30-day minimum retention for safety monitoring.

data residency

Review of enterprise documentation and privacy policies

Evidence
OpenAI Enterprise Documentation โ€” US-based processing, enterprise options for data residency
highVerified: 2026-07-09
training data optout

Analysis of privacy policy and data usage terms

Evidence
OpenAI Privacy Policy โ€” No training on API data by default
highVerified: 2026-07-09
data retention

Review of terms of service and data retention policies

Evidence
OpenAI Data Usage Policy โ€” 30-day retention for safety monitoring, deletable after
highVerified: 2026-07-09
pii handling

Review of data protection capabilities and customer responsibilities

Evidence
OpenAI Privacy Documentation โ€” Customer responsible for PII handling
mediumVerified: 2026-07-09
compliance certifications

Verification of compliance certifications and audit reports

Evidence
OpenAI Trust Portal โ€” SOC 2 Type II, GDPR compliant
highVerified: 2026-07-09
zero data retention

Review of data handling practices

Evidence
OpenAI Enterprise Options โ€” Minimum 30-day retention for safety, no true zero retention
mediumVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+

Excellent explainability via chain-of-thought reasoning. Transparent problem-solving process visible to users.

explainability

Evaluation of reasoning transparency and explanation capabilities

Evidence
Chain-of-Thought Reasoning โ€” Extended chain-of-thought visible to users
highVerified: 2026-07-09
hallucination rate

Testing on factual QA datasets and real-world usage

Evidence
OpenAI Benchmarks โ€” Reduced hallucination via chain-of-thought verification
highVerified: 2026-07-09
bias fairness

Evaluation on bias benchmarks and diverse demographic testing

Evidence
OpenAI Safety Research โ€” Ongoing bias testing and mitigation
mediumVerified: 2026-07-09
uncertainty quantification

Assessment of confidence expression in outputs

Evidence
Model Behavior โ€” Good uncertainty expression through reasoning process
highVerified: 2026-07-09
model card quality

Review of documentation completeness and clarity

Evidence
OpenAI Model Documentation โ€” Comprehensive model documentation
highVerified: 2026-07-09
training data transparency

Review of public disclosures about training data

Evidence
OpenAI Public Statements โ€” General description of training approach
mediumVerified: 2026-07-09
guardrails

Analysis of built-in safety mechanisms

Evidence
OpenAI Safety Features โ€” Comprehensive safety guardrails
highVerified: 2026-07-09
โš™๏ธOperational Excellence
+

Deprecated: API shutdown scheduled 2026-10-23, migration target GPT-5.5. Versioning and ecosystem scores reduced to reflect deprecation.

api design quality

Review of API design, consistency, and feature completeness

Evidence
OpenAI API Documentation โ€” Well-designed RESTful API
highVerified: 2026-07-09
sdk quality

Review of SDK quality, documentation, and maintenance

Evidence
OpenAI SDKs โ€” Official SDKs for Python, Node.js, actively maintained
highVerified: 2026-07-09
versioning policy

Review of versioning policy and historical practices

Evidence
OpenAI API Versioning โ€” Clear versioning with deprecation notices
OpenAI Deprecations โ€” o1 (o1-2024-12-17) API shutdown 2026-10-23, replacement gpt-5.5; o1-pro shutdown 2026-10-23, replacement gpt-5.5-pro; o1-preview shut down 2025-07-28 and o1-mini 2025-10-27
highVerified: 2026-07-09
monitoring observability

Review of available monitoring tools and metrics

Evidence
OpenAI Platform โ€” Comprehensive usage dashboard
highVerified: 2026-07-09
support quality

Assessment of documentation, community, and support responsiveness

Evidence
OpenAI Support โ€” Comprehensive support with enterprise SLAs
highVerified: 2026-07-09
ecosystem maturity

Analysis of third-party integrations and tools

Evidence
OpenAI Ecosystem โ€” Mature ecosystem with extensive integrations
highVerified: 2026-07-09
license terms

Review of licensing terms and restrictions

Evidence
OpenAI Terms of Service โ€” Standard commercial terms, enterprise agreements available
highVerified: 2026-07-09
Strengths
  • +Best-in-class reasoning with 78.3% GPQA Diamond
  • +Visible chain-of-thought for transparent problem-solving
  • +Exceptional mathematical capabilities (83% on AIME)
  • +Strong coding performance (57.1% SWE-bench)
  • +Excellent for complex analytical and research tasks
  • +High explainability via reasoning traces
Limitations
  • !High latency (4.5s p50, 8.2s p95) due to reasoning overhead
  • !Not suitable for real-time applications
  • !30-day minimum data retention (not ephemeral)
  • !Not HIPAA eligible
  • !Higher cost due to extended reasoning compute
  • !Reasoning overhead may be unnecessary for simple tasks
  • !DEPRECATED: removed from ChatGPT; o1 variants and o1-pro API shutdown 2026-10-23 โ€” migrate to GPT-5.5
Metadata
pricing
input: $15.00 per 1M tokens
output: $60.00 per 1M tokens
notes: Premium reasoning model pricing, significantly higher than standard models. Pricing applies until API shutdown 2026-10-23.
last verified: 2026-07-09
context window: 128000
languages
0: English
1: Spanish
2: French
3: German
4: Italian
5: Portuguese
6: Japanese
7: Korean
8: Chinese
modalities
0: text
api endpoint: https://api.openai.com/v1/chat/completions
open source: false
architecture: Transformer-based with extended chain-of-thought reasoning
parameters: Not disclosed

Use Case Ratings

code generation

Excellent coding with 57.1% SWE-bench and 79.2% HumanEval. Chain-of-thought helps with complex algorithms.

customer support

Good capabilities but high latency (4.5s) may impact customer experience. Better for complex issues.

content creation

Good content generation but reasoning focus may add unnecessary latency for creative tasks.

data analysis

Exceptional analytical capabilities with chain-of-thought reasoning. Best for complex analysis.

research assistant

Outstanding research capabilities with transparent reasoning. Excellent for complex research tasks.

legal compliance

Good reasoning for legal analysis but 30-day retention may be concern for some use cases.

healthcare

Good reasoning but not HIPAA eligible. 30-day retention may be concern for healthcare data.

financial analysis

Outstanding for complex financial modeling and analysis with transparent reasoning.

education

Exceptional for education with visible chain-of-thought. Perfect for teaching problem-solving.

creative writing

Competent but reasoning focus may reduce creative spontaneity. Higher latency for creative tasks.