What is GPT 6.1 Sol?
GPT 6.1 Sol is OpenAI's upgrade to GPT 6 Sol, released on September 29, 2026 at DevDay. It runs on Unifically as openai/gpt-6.1-sol. OpenAI built it to get close to GPT 6 Astra on agentic coding, computer use, and professional work at one fifth of Astra's input and output token prices. It keeps GPT 6 Sol's per-token price and halves the cached input price. It takes text and image input, returns text, and works across a 1,050,000-token context window with a 128,000-token max output and knowledge up to April 30, 2026. Reasoning effort runs from low up to max.
What GPT 6.1 Sol can do
Astra-level coding for a fifth of the price
On DeepSWE v1.1, which runs long software tasks in real codebases, GPT 6.1 Sol scores 75.2%. That is level with GPT 6 Astra at 74.1% and 6.4 points above GPT 6 Sol's best of 68.8%, reached at a lower reasoning effort than GPT 6 Sol needed.
Computer use within two points of Astra
OSWorld 2.0 has agents work through long desktop and browser workflows. At max effort GPT 6.1 Sol scores 71.4% on the offline set, 7 points above GPT 6 Sol at less than half its cost, and 2.1 points behind GPT 6 Astra at about one seventh of Astra's cost per task.
Science workflows at a quarter of the cost
Terminal-Bench Science 0.1 has agents analyze data, run simulations, and prove theorems from a terminal. GPT 6.1 Sol solves 57.0% at max effort for $5.47 a task, more than double GPT 6 Sol's score. Claude Opus 5.5 costs $23.21 a task there and GPT 6 Astra $23.80.
Fewer factual errors
On OpenAI's set of hard factual prompts, GPT 6.1 Sol at low effort answers with an error 7.7% of the time against 11.4% for GPT 6 Sol, and stays within 1.9 points of Astra at every effort. The independent Artificial Analysis Intelligence Index scores it 52 at max effort, 10th of 221 models.
Best for
Coding agents
Repo-scale changes, test-and-fix loops, and code review. 75.2% on DeepSWE v1.1.
Computer use
Browser and desktop workflows. 71.4% on OSWorld 2.0 at max effort.
Business workflows
Multi-step processes across apps. 2.2 points above Claude Opus 5.5 on AutomationBench at medium effort, at about a third of the cost.
Document analysis
Tables, charts, and fine print in long PDFs. Ahead of Claude Opus 5.5 on GDP.pdf at less than half the cost per task.
Scientific computing
Data analysis and simulation from the command line for $5.47 a Terminal-Bench Science task.
Cached agent loops
Large repos or tool lists re-read on every step, with cached input at 5% of the input rate.
Use cases
Swap GPT 6.1 Sol in as the default model behind a coding agent that plans a change, edits across the repo, runs the tests, and reviews its own diff. Keep the repository and tool definitions in a stable prefix so each step re-reads them at the cache-read rate. Point it at a browser or desktop for form filling, QA passes, and research runs. Load contracts, filings, or financial reports into the 1M-token window and ask for answers that cite the table or chart they came from. Teams already paying for Astra can move most agent steps to 6.1 Sol and keep Astra for the hardest science and research tasks.
Limitations
Astra is still the stronger model at the top end. It scores 68.1% on Terminal-Bench Science 0.1 against 57.0%, and 73.5% on OSWorld 2.0 against 71.4%. Keep Astra for the most difficult scientific research.
It is not fast at high effort. Artificial Analysis measures about 67 output tokens per second at max effort, below the average for its price range, and a long thinking phase before the first answer token. Use low or medium effort for interactive work.
Reasoning cannot be switched off. The none and minimal efforts that GPT 6 Sol accepts are not supported, so the cheapest request still reasons at low.
Long prompts cost more. Once a request's prompt passes 272,000 tokens, the whole request bills at the long-context rate: double on input, cache read, and cache write, and 1.5x on output.
GPT 6.1 Sol vs GPT 6 Sol
The price per input and output token is the same, so there is little reason to stay on GPT 6 Sol. GPT 6.1 Sol scores higher on every benchmark OpenAI published: DeepSWE v1.1 (75.2% against 68.8%), OSWorld 2.0 (71.4% against 64.4%), AutomationBench (4.8 points higher at medium effort), and Terminal-Bench Science (more than double). Cached input costs half as much, and it reports broken tools to the user more often, failing to disclose one in 2.1% of tests against 4.9%. Keep GPT 6 Sol only if you rely on reasoning_effort: "none".
When to use GPT 6.1 Sol
Use GPT 6.1 Sol as the everyday model for coding agents, computer use, and document-heavy analysis, where you want close to Astra's results for a fifth of the token price. Step up to GPT 6 Astra for the hardest science and research tasks, and drop to GPT 6 Luna for high-volume classification, extraction, and short replies.
API examples
Call GPT 6.1 Sol from any language by POSTing to /v1/chat/completions, the OpenAI-compatible endpoint shared by every language model on the platform. Full parameter docs live at docs.unifically.com/models/llm/openai/gpt-6.1-sol.
curl -X POST https://api.unifically.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "openai/gpt-6.1-sol",
"reasoning_effort": "medium",
"messages": [
{ "role": "user", "content": "Review this pull request and list any bugs you find, most serious first." }
]
}'
The response comes back synchronously with the completion. Set "stream": true to receive tokens as they generate, and raise reasoning_effort to high or max for hard multi-step tasks.
FAQs
People also ask
openai/gpt-6.1-sol, called through the OpenAI-compatible POST /v1/chat/completions endpoint with your Unifically API key. The same ID also works on /v1/responses and /v1/messages, and tool calls work on all three.
1,050,000 tokens, with a max output of 128,000 tokens and a knowledge cutoff of April 30, 2026. It takes text and image input and returns text.
It is the same price per token with better results. It scores 75.2% on DeepSWE v1.1 against 68.8%, 71.4% on OSWorld 2.0 against 64.4%, and more than doubles GPT 6 Sol on Terminal-Bench Science 0.1. Cached input costs half as much, and at low effort it makes factual errors on 7.7% of hard prompts against 11.4%.
Close on coding and computer use, not across the board. It edges Astra on DeepSWE v1.1, 75.2% against 74.1%, and lands 2.1 points behind it on OSWorld 2.0. Astra still leads on Terminal-Bench Science 0.1, 68.1% against 57.0%, but costs $23.80 a task there against $5.47.
low, medium (the default), high, xhigh, and max, set with reasoning_effort. OpenAI dropped none and minimal for this model, so every request does some reasoning.
Billing switches to the long-context rate for the whole request once the prompt passes 272,000 tokens. Input, cache read, and cache write bill at double the standard rate and output at 1.5x. Both rates are listed on the pricing page.
Not yet. OpenAI announced an Ultrafast option with up to 8x faster token generation, rolling out to Codex first. This page covers the standard openai/gpt-6.1-sol model.
