GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท gpt-4-1-nano

GPT-4.1 nano

vgpt-4.1-nano-2025-04-14

OpenAI

Modeldeprecatedefficientlow-latencycost-effective
80
Strong
About This Model

DEPRECATED: OpenAI announced 2026-04-22 that gpt-4.1-nano's API shuts down 2026-10-23; recommended replacement is gpt-5.4-nano. Historically OpenAI's smallest and most efficient GPT-4.1 variant for high-volume, cost-sensitive applications, with a 1,047,576-token context window.

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+

Basic performance optimized for speed and efficiency. Best for simple tasks where ultra-low latency and cost are priorities.

task accuracy code

Industry-standard coding benchmarks measuring basic programming tasks

Evidence
HumanEval Benchmark โ€” 29.4% pass rate
highVerified: 2026-07-09
task accuracy reasoning

Basic reasoning benchmarks

Evidence
MATH Benchmark โ€” 35% on mathematical reasoning tasks
mediumVerified: 2026-07-09
task accuracy general

Crowdsourced comparisons and knowledge testing

Evidence
MMLU Benchmark โ€” 50.3% on multitask language understanding
LMSYS Chatbot Arena โ€” 1050 ELO (Entry-level performance)
highVerified: 2026-07-09
output consistency

Internal testing with repeated prompts

Evidence
OpenAI Internal Testing โ€” Reasonable consistency for simple tasks
mediumVerified: 2026-07-09
latency p50

Median latency for API requests

Evidence
OpenAI Documentation โ€” Ultra-fast response time ~0.4s
highVerified: 2026-07-09
latency p95

95th percentile response time

Evidence
Community benchmarking โ€” p95 latency ~0.8s
highVerified: 2026-07-09
context window

Official specification from provider

Evidence
OpenAI Model Page: gpt-4.1-nano โ€” 1,047,576 token context window; 32,768 max output tokens
highVerified: 2026-07-09
uptime

Historical uptime data from official status page

Evidence
OpenAI Status Page โ€” 99.9% uptime (last 90 days)
highVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+

Good security posture with standard OpenAI safety measures. Smaller model may have slightly lower resistance to adversarial attacks.

prompt injection resistance

Testing against OWASP LLM01 prompt injection attacks

Evidence
OpenAI Safety Testing โ€” Moderate resistance to prompt injection
mediumVerified: 2026-07-09
jailbreak resistance

Testing against adversarial prompt datasets

Evidence
OpenAI Safety Evaluations โ€” Basic safety mechanisms in place
mediumVerified: 2026-07-09
data leakage prevention

Analysis of privacy policies and data handling practices

Evidence
OpenAI Privacy Policy โ€” API data not used for training by default
mediumVerified: 2026-07-09
output safety

Safety testing across harmful content categories

Evidence
OpenAI Safety Benchmarks โ€” Standard content filtering applied
highVerified: 2026-07-09
api security

Review of API security features and best practices

Evidence
OpenAI API Documentation โ€” API key authentication, HTTPS only, rate limiting
highVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+

Standard OpenAI privacy practices. 30-day data retention for abuse monitoring.

data residency

Review of enterprise documentation and privacy policies

Evidence
OpenAI Documentation โ€” US-based infrastructure
highVerified: 2026-07-09
training data optout

Analysis of privacy policy and data usage terms

Evidence
OpenAI Privacy Policy โ€” API data not used for training by default
highVerified: 2026-07-09
data retention

Review of terms of service and data retention policies

Evidence
OpenAI Terms of Service โ€” API data retained for 30 days for abuse monitoring
highVerified: 2026-07-09
pii handling

Review of data protection capabilities

Evidence
OpenAI Privacy Documentation โ€” Customer responsible for PII redaction
mediumVerified: 2026-07-09
compliance certifications

Verification of compliance certifications

Evidence
OpenAI Trust Portal โ€” SOC 2 Type II, GDPR compliant
highVerified: 2026-07-09
zero data retention

Review of data handling practices

Evidence
OpenAI API Documentation โ€” 30-day retention for abuse monitoring
highVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+

Basic transparency features. Smaller model size limits explainability depth. Higher hallucination rate than premium models.

explainability

Evaluation of reasoning transparency

Evidence
Model Behavior โ€” Basic explanations, less detailed than larger models
mediumVerified: 2026-07-09
hallucination rate

Testing on factual QA datasets

Evidence
SimpleQA Benchmark โ€” Moderate hallucination rate on simple queries
mediumVerified: 2026-07-09
bias fairness

Evaluation on bias benchmarks

Evidence
OpenAI Safety Report โ€” Regular bias testing applied
mediumVerified: 2026-07-09
uncertainty quantification

Qualitative assessment of confidence expression

Evidence
Model Behavior โ€” Limited uncertainty expression
mediumVerified: 2026-07-09
model card quality

Review of documentation completeness

Evidence
OpenAI Model Documentation โ€” Good documentation with capabilities and limitations
highVerified: 2026-07-09
training data transparency

Review of public disclosures about training data

Evidence
OpenAI Public Statements โ€” General description provided
mediumVerified: 2026-07-09
guardrails

Analysis of built-in safety mechanisms

Evidence
OpenAI Safety Systems โ€” Standard safety guardrails
highVerified: 2026-07-09
โš™๏ธOperational Excellence
+

Excellent operational maturity leveraging OpenAI's established infrastructure. Versioning score reduced to reflect deprecation: API shutdown 2026-10-23, migrate to gpt-5.4-nano.

api design quality

Review of API design and consistency

Evidence
OpenAI API Documentation โ€” Consistent RESTful API across model family
highVerified: 2026-07-09
sdk quality

Review of SDK quality and maintenance

Evidence
OpenAI SDKs โ€” Official SDKs for Python, Node.js
highVerified: 2026-07-09
versioning policy

Review of versioning policy

Evidence
OpenAI API Versioning โ€” Clear versioning with deprecation notices
OpenAI Deprecations โ€” gpt-4.1-nano deprecation announced 2026-04-22; API shutdown 2026-10-23; recommended replacement gpt-5.4-nano
highVerified: 2026-07-09
monitoring observability

Review of monitoring tools

Evidence
OpenAI Dashboard โ€” Usage dashboard with basic metrics
mediumVerified: 2026-07-09
support quality

Assessment of support channels

Evidence
OpenAI Support โ€” Email support, forum community
highVerified: 2026-07-09
ecosystem maturity

Analysis of third-party integrations

Evidence
GitHub Ecosystem โ€” Mature ecosystem with extensive integrations
highVerified: 2026-07-09
license terms

Review of licensing terms

Evidence
OpenAI Terms of Service โ€” Standard commercial terms
highVerified: 2026-07-09
Strengths
  • +Ultra-low latency (~0.4s p50) ideal for real-time applications
  • +Most cost-effective option in GPT-4.1 family ($0.10/$0.40 per 1M)
  • +Good for high-volume, simple tasks
  • +Very large context window (1,047,576 tokens) at nano pricing
  • +Same API and ecosystem as premium OpenAI models
  • +Reliable uptime and infrastructure
Limitations
  • !Limited coding capabilities (29.4% HumanEval)
  • !Basic reasoning and knowledge (50.3% MMLU)
  • !Higher hallucination rate than larger models
  • !Not suitable for complex or specialized tasks
  • !30-day data retention
  • !DEPRECATED: API shutdown 2026-10-23 (announced 2026-04-22); migrate to gpt-5.4-nano
Metadata
pricing
input: $0.10 per 1M tokens
output: $0.40 per 1M tokens
notes: Cached input $0.025 per 1M. Confirmed on official model page 2026-07-09; model deprecated with API shutdown 2026-10-23.
last verified: 2026-07-09
context window: 1047576
max output: 32768
languages
0: English
1: Spanish
2: French
3: German
4: Italian
5: Portuguese
6: Japanese
7: Korean
8: Chinese
modalities
0: text
api endpoint: https://api.openai.com/v1/chat/completions
open source: false
architecture: Transformer-based, optimized for efficiency
parameters: Not disclosed (small)

Use Case Ratings

code generation

Basic code generation for simple tasks. 29.4% HumanEval indicates limited capability for complex programming.

customer support

Good for high-volume, simple customer queries. Fast response times make it suitable for basic support automation.

content creation

Adequate for simple content tasks. Limited creativity and depth compared to larger models.

data analysis

Basic data interpretation. Not suitable for complex analytical tasks.

research assistant

Suitable for simple research queries and summaries. Limited depth for complex topics.

legal compliance

Not recommended for legal applications due to limited accuracy and reasoning.

healthcare

Not suitable for healthcare applications. Lacks accuracy and HIPAA eligibility.

financial analysis

Basic financial calculations only. Not suitable for complex financial modeling.

education

Suitable for basic educational content and simple tutoring. Limited for advanced topics.

creative writing

Basic creative writing. Less nuanced and creative than larger models.