On the three benchmarks OpenAI published for its new financial-services product, its latest model beat the previous one by 3 points, by 9.7 points, and by 34 points. The 34-point jump wasn't in financial reasoning or in finding numbers inside documents. It was in making slides.
That lopsidedness is the most honest thing in the announcement, and it tells you what ChatGPT for Financial Services is actually selling.
Now available: ChatGPT for Financial Services.
— OpenAI (@OpenAI) September 10, 2026
This is a tailored ChatGPT Work experience that combines built-in financial data with GPT-6 Astra’s reasoning.
Teams can develop research, build financial models, and create customized client materials.https://t.co/6WP5OJdnE8 pic.twitter.com/AundGG3jtc
What OpenAI actually shipped
ChatGPT for Financial Services is a version of ChatGPT Work aimed at investment banking and equity research, powered by GPT-6 Astra. It was shaped with Morgan Stanley and Evercore as design partners, and it arrives with premium financial data already inside it — earnings transcripts, financial statements, company fundamentals and private-company records from Daloopa, PitchBook, LSEG News and Crunchbase.
The named workflows are exactly the ones that fill an analyst's week: value analysis, LBO modelling, buyer screening, earnings analysis and pitchbook preparation. Administrators can publish the firm's own Excel, Word and PowerPoint templates so output lands in house format rather than generic AI formatting.
There's a catch buried in the benchmark numbers, though, and it's worth holding onto until the end.
The number that jumped 34 points
Look at how unevenly the model improved across the three capabilities OpenAI says this work depends on.

The model got a little better at reading documents, meaningfully better at finding facts, and radically better at producing the deliverable. (Chart: BougainWell · Data: OpenAI)
In relative terms the gap is even starker. BoxBench, which tests reasoning over documents in business workflows, improved about 4%. OfficeQA Pro, which tests finding and analysing information across US Treasury Bulletins including tables and footnotes, improved about 16%. The slide-creation win rate improved about 157% — it went from winning roughly one head-to-head in five to winning about five in nine.
Read that as a sentence rather than a chart: the model did not mainly get smarter about finance. It got much better at turning thinking into an artifact someone can hand to a client.
That is a genuinely useful thing to be better at, because the artifact is where junior finance work actually goes. But it is a different claim from "the AI now understands markets." Document reasoning at 77% versus 74% looks like a capability approaching a plateau, not one breaking open.
The real unlock is procurement, not intelligence
The line in OpenAI's announcement that will change more workdays than any benchmark is this one: teams can start working with these datasets immediately, with no separate contracts to negotiate or connectors to set up.
Anyone who has tried to get a data feed approved inside a regulated institution knows that sentence describes months of work, not minutes. Vendor review, entitlement mapping, security sign-off, a connector that breaks quietly — that pipeline is the actual reason AI pilots stall at big banks. It is rarely the model.
So the interesting move here isn't that OpenAI made a finance-flavoured chatbot. It's that OpenAI licensed the data, indexed it, and is hosting it itself — becoming a data distributor as well as a model vendor. Granular citations fall out of that decision: because OpenAI holds the index, it can trace a figure back to its source document instead of hoping a third-party API returns something checkable.
Three tiers of closeness
Read the partner list carefully and you'll see it isn't one list. It's three, and the tier a provider sits in tells you how much leverage it kept.
| Tier | What it means | Who's there |
|---|---|---|
| Hosted inside | OpenAI indexes and serves the data directly | Daloopa, PitchBook, LSEG News, Crunchbase |
| Sign-in passthrough | Your existing subscription is recognised via ChatGPT login | S&P Capital IQ, LSEG, MSCI, Dow Jones Factiva, Moody's |
| Connector | Reachable, but you wire it up | Datasite, Box, Preqin, FactSet, Intapp (50+ total) |
Tier one is shelf space at the front of the store. Tier three is being in the warehouse. LSEG appears in both tier one and tier two, which is the hedge you'd expect from a provider that wants the distribution without handing over the whole relationship.
The trade for tier-one providers is real: shelf space now, in exchange for the risk that the end user stops thinking of you as a product and starts thinking of you as a field in someone else's answer. Data businesses that become invisible plumbing historically lose pricing power, even as volume rises.
Read the fine print
Here's the catch promised earlier. Two of the three headline numbers need an asterisk.
BoxBench is, by OpenAI's own note, an evaluation provided by an external partner rather than one OpenAI built and controls. And the slide-creation figure is not a quality score at all — it's a win rate against Opus 5 in human comparisons. A win rate can triple without the underlying output getting three times better; it only means human judges preferred it more often in a head-to-head. Those are useful signals, not the same as an absolute measure.
Availability comes with its own gate. The product is offered to eligible financial institutions through a sales conversation, with no public pricing and no self-serve sign-up. Whatever this costs, it is priced per firm, behind a conversation — which is its own signal about who OpenAI thinks the buyer is.
On the governance side the list is genuinely enterprise-grade: SAML single sign-on, SCIM provisioning, role-based access control, business data excluded from training by default, encryption at rest and in transit, configurable retention, workspace log export for compliance teams, and multiple workspaces so firms can enforce information barriers between deal teams. For a bank, that last item isn't a feature — it's the entry ticket.
What to Take Away
- When a vendor publishes several benchmarks, read the spread, not the average. Gains of +3, +9.7 and +34 across three tests describe a product far more precisely than any single headline number would.
- A "win rate" is not a quality score. It tells you how often judges preferred one output to another, which can move dramatically without the work getting dramatically better. Always ask "won against what, judged by whom."
- The bottleneck AI removes at large institutions is usually procurement, not intelligence. Bundling licensed data so there are no new contracts to sign is a bigger unlock than a few benchmark points.
- Watch which tier a data provider accepted. Hosted-and-indexed buys distribution but risks becoming invisible plumbing; passthrough keeps the customer relationship. The tier is the strategy.
- The work most exposed here is artifact production — the deck, the model, the formatted note — not judgement about what the numbers mean. That distinction is worth remembering whichever side of it your job sits on.
Featured image: OpenAI. Chart: BougainWell, built from OpenAI's published benchmark figures. This article is for general information only and is not investment advice.
Sources
- OpenAI (@OpenAI) on X — announcing ChatGPT for Financial Services
- OpenAI — Introducing ChatGPT for Financial Services (September 10, 2026)
All analysis and opinions in this article are BougainWell's own.



