Every AI vendor that sells to banks and funds now has a deployment slide: multi-tenant SaaS, single tenant, private cloud, your VPC, on-premise. Buyers ask for it because their own annual reports now explain why. We read the 10-Ks of every bank, broker, asset manager and insurer that filed one with the SEC between January and September 2026, 550 companies, and measured what they say about AI. 84% mention artificial intelligence, up from 14% of the same group in 2023. Of the 384 that discuss AI in their risk factors, 54% name the risk of relying on third-party AI vendors and 37% the risk of confidential data leaking through AI tools. Those are the two risks a deployment model controls. The rest of the list, from wrong answers to AI-enabled cyberattacks, it doesn't.
TLDR:
- 462 of 550 financial-sector 10-K filers (84%) mentioned AI in 2026, against 59% of all 10-K filers. 239 (43%) mentioned generative AI; in 2023, none of them did.
- 384 discuss AI inside Item 1A Risk Factors. 60% of those flag third-party AI vendors, confidential-data exposure, or both. Among brokers, investment banks and asset managers it is 74%.
- On-premise, customer-VPC and single-tenant SaaS answer those two risks to different degrees. None of them fixes inaccurate output, which 61% of the filings also flag; that depends on the data the model reads. A vendor checklist is below.
AI is now in almost every financial 10-K
We counted every company that filed a 10-K between January 1 and September 30 of each year, using the SEC's own EDGAR indexes, and checked each filing for the phrase "artificial intelligence" with EDGAR full-text search. Banks are SIC codes 60xx, brokers, dealers, exchanges and asset managers 62xx, insurers 63xx–64xx. The share mentioning AI rose every year in every group.
| 10-K filers mentioning AI (filers in 2026) | 2023 | 2026 | Generative AI, 2026 |
|---|---|---|---|
| Banks (332) | 8.9% | 81.0% | 43.1% |
| Brokers & asset managers (90) | 17.2% | 87.8% | 50.0% |
| Insurers (128) | 27.1% | 89.1% | 39.8% |
| All three (550) | 14.3% | 84.0% | 43.5% |
| All 10-K filers, every sector (6,245) | 16.1% | 59.0% | 25.0% |
Source: FocusAlpha analysis of SEC EDGAR full-text search and EDGAR form indexes, 10-Ks filed January 1 to September 30 of each year. For all three groups combined, the share was 46.3% in 2024 and 67.4% in 2025. Methodology below.
Figure 1: Share of 10-K filers mentioning "artificial intelligence", by filing year. Source: FocusAlpha analysis of SEC EDGAR full-text search.
In 2023 the financial sector sat slightly below the market average. Three years later it is 25 points above it. Banks moved fastest, from 9% to 81%, and generative AI went from absent to 43% of financial filers in the same period. Brokers and asset managers are the most likely to name it specifically: half of them do.
What the risk factors actually say
A mention can be a product announcement. A risk factor is a legal disclosure of what could go wrong. For the 462 companies that mentioned AI in 2026, we parsed the main 10-K document, located Item 1A Risk Factors, and tagged every sentence that mentions AI with keyword rules for six risks. 384 companies discuss AI inside their risk factors. This is what they worry about:
Figure 2: Risk themes in AI sentences in Item 1A, 10-Ks filed January to September 2026. Source: FocusAlpha analysis of 10-K text from SEC EDGAR.
The two dark bars are the risks that a deployment decision changes. The sentences behind them are specific. Goldman Sachs, in its 2025 10-K, says:
"we rely on AI models developed by third parties, and, to that extent, are dependent in part on the manner in which those third parties develop and train their models ... matters over which we may have limited visibility."
Capital One describes the leak path in one sentence in its 10-K: if employees or service providers use third-party AI tools, it "may lead to the inadvertent or unauthorized disclosure or incorporation of our sensitive and confidential information, including personal information, into third-party systems or publicly available or third-party training sets." Lazard, an advisory firm, states the input problem plainly in its 10-K: "AI tools may require inputting or processing sensitive information, including proprietary Lazard information, as well as confidential client and third party data." JPMorganChase adds the agentic version in its 10-K, warning of data loss if AI systems, "particularly agentic systems", lack safeguards "to prevent systems from accessing sensitive data sources."
Much of this text is shared: 52 filings warn that AI could lead to "the release of" private, personal or confidential information, often in near-identical sentences. Few describe the control. Stifel is a rare one, saying in its 10-K that it mitigates AI privacy risk "by relying on proprietary or 'walled-garden' environments to enhance data protection and operational controls and maintain confidentiality."
The concern is sharpest at the firms that buy research tools.
| Companies discussing AI in Item 1A, 2026 | Count | Third-party AI vendors | Confidential-data exposure | Either |
|---|---|---|---|---|
| Brokers & asset managers | 69 | 72% | 48% | 74% |
| Banks | 221 | 56% | 37% | 60% |
| Insurers | 94 | 38% | 30% | 48% |
| All three | 384 | 54% | 37% | 60% |
Source: FocusAlpha analysis of 10-K text from SEC EDGAR. A company counts once per theme.
On-premise, private cloud, VPC, single tenant: what each one changes
These terms get used loosely, so here is what each one means in practice. The question that separates them is where your prompts, documents and outputs physically go, and who controls that path.
- Multi-tenant SaaS. The software runs in the vendor's cloud, shared with other customers, and calls the vendor's model provider. Vendor risk is managed by contract alone. Your data leaves your network and is separated from other customers' data only logically.
- Single-tenant SaaS. A dedicated instance in the vendor's cloud, usually still calling the vendor's model provider. Same contract, smaller blast radius. Your data still leaves your network.
- SaaS with your keys or storage (BYOK, BYOB). The vendor's cloud and model provider, but stored data sits under your encryption keys or in your own storage bucket, and you can revoke the keys. Processing is still the vendor's.
- Customer VPC ("private cloud"). The vendor's software runs in your AWS, Azure or Google Cloud account and calls a model endpoint you choose, often in that same account. You see and control every outbound connection, and prompts and documents stay in your account if the model does too.
- On-premise or air-gapped. Your data center, and models you host on your own hardware. The lowest exposure on both risks, at the highest cost.
Two details decide whether a "private" deployment is private.
First, ask where inference runs. An application in your VPC that sends every prompt to a model API on the public internet has moved the user interface, not the data. The private version calls a model endpoint inside your cloud perimeter, such as the frontier models offered through Amazon Bedrock, Google Cloud Vertex AI or Microsoft's Azure AI platform. Fully on-premise deployments are limited to models you can run on your own hardware: open-weight models such as OpenAI's gpt-oss, or the few proprietary models licensed for on-premise use. Most frontier models are sold as cloud endpoints.
Second, the questions are sensitive too. A fund's question log shows what it is researching, and that is often close to what it holds or plans to trade. CNL Strategic Capital, which invests in private companies, names the failure mode in its 10-K: "a user may input confidential information, including material non-public information or personal identifiable information, into artificial intelligence technologies," where it can become accessible to "third-party artificial intelligence applications and users, including competitors." Even a query about public data can reveal intent, so the path your agent takes to fetch outside data belongs in the security review.
What vendors say they offer
We read the security and deployment pages of the finance AI vendors buyers most often compare, on October 4, 2026. Vendors often agree to more in a contract than they publish, so treat this as a starting point for questions, not a ranking.
| Vendor | Deployment options stated | Certifications | Trains on your data? |
|---|---|---|---|
| Rogo | Single tenant; data in "siloed environments" | SOC 2, ISO 27001 | No, per its security page |
| Hebbia | Dedicated tenant "if needed"; US or EU processing | SOC 2 Type 2, ISO 27001, ISO 42001 | No, nor its model providers |
| AlphaSense | SaaS, with your own keys (BYOK) or S3 bucket (BYOB) | SOC 2 Type 2, ISO 27001 | No, and zero-retention LLM providers |
| Unique AI | Multi-tenant, single tenant, self-hosted in your cloud, or on-premise | SOC 2, ISO | Not stated |
| FocusAlpha | SaaS; the Fund plan can be deployed in your VPC | SOC 2 Type II | — |
The finance-native research platforms mostly stop at single tenant or customer-managed keys. Customer-VPC and on-premise options are more common from platform vendors, so if your policy requires either, ask for it in the first call, not the last.
What regulators ask for
None of the documents below requires on-premise AI. They require that a firm keep responsibility for its vendors, know where its data goes, and be able to show it.
- FINRA Regulatory Notice 24-09 (June 27, 2024): FINRA rules apply "whether member firms are directly developing Gen AI tools for their proprietary use or when leveraging the technology of a third party." Its 2026 oversight report lists, as an effective practice, contract language "that prohibits firm or customer sensitive information from being ingested into a third-party vendor's open-source GenAI tool."
- SEC FY2026 examination priorities: exams will "assess whether firms have implemented adequate policies and procedures to monitor and/or supervise their use of AI technologies," and Regulation S-P reviews will focus on "oversight of third-party vendors."
- SEC Regulation S-P, as amended in 2024: service providers must notify a covered institution "as soon as possible, but no later than 72 hours" after a breach. Larger entities had to comply by December 3, 2025, smaller ones by June 3, 2026.
- NYDFS guidance on third-party service providers (October 21, 2025): covered entities "may not delegate responsibility" to a provider; contracts should require it "to disclose where data may be stored, processed, or accessed," and firms should consider a clause on "the acceptable use of Artificial Intelligence" and "whether the Covered Entity's data may be used to train AI models."
- Federal Reserve, OCC and FDIC SR 26-2 (April 17, 2026), the model risk guidance that replaced SR 11-7: "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." Banks' own risk management and governance practices are expected to fill the gap.
- EU DORA (Regulation 2022/2554, applying since January 17, 2025): financial entities "remain fully responsible" when they use ICT providers (Article 28), and contracts must state "where data is to be processed, including the storage location" (Article 30).
The pattern is consistent. A deployment model is a way to answer "where does our data go, and who touches it?" with evidence. A customer-VPC deployment answers it with your own network logs. A SaaS deployment answers it with the vendor's SOC 2 report, subprocessor list and contract.
A checklist for any AI vendor, Rogo-style platforms included
- Where does inference run? Which model, which version, in whose cloud account and region?
- What leaves our network? Prompts, uploaded documents, embeddings, logs, telemetry, support access.
- Retention. How long do the vendor and its model provider keep prompts and outputs? Is zero data retention available?
- Training. Is training on our data excluded in the contract, for the vendor and every model provider it uses?
- Subprocessors. The full list, and advance notice of changes.
- Data location. Where data is stored, processed and accessed, as NYDFS and DORA expect contracts to state.
- Assurance. A SOC 2 Type II report whose scope covers the AI product, its review period, and a bridge letter for the gap since.
- Incident notice. A commitment to notify you of a breach within 72 hours, which Regulation S-P requires your vendor oversight to include.
- Isolation and keys. Single tenant or shared, and whether you hold the encryption keys.
- Outside data. When the agent needs filings, transcripts, market data or the web, where do those calls go, and what do the queries reveal?
- Citations. Can every number in an answer be traced to the source document?
Private deployment doesn't fix wrong answers
The third-largest risk in Figure 2 is inaccurate or biased output, flagged by 61% of the filings. Moving the model into your VPC does nothing for it. An agent behind your firewall still needs public-company data, and it answers only as well as that data allows. If it pulls filings from the open web, it inherits parsing errors, and every lookup leaves your perimeter.
This analysis is an example of the work involved. Most of it was not reading 10-Ks. It was finding the right 550 filers, choosing one original filing per company over its amendments, and cutting Item 1A out of HTML that every company formats differently; we located it in 405 of 462 filings. That is the layer FocusAlpha provides: SEC filings, earnings-call transcripts and ownership data, structured, with each value cited back to its source document. Splitting a 10-K into its Items is one request to the filing items endpoint, and the same data reaches Claude through the FocusAlpha MCP server:
GET /v1/filings/items?ticker=GS
&filing_type=10-K&year=2025&item=1A
On the Fund plan, FocusAlpha can be deployed inside your own VPC, next to your agents, and we are SOC 2 Type II. Starter and Professional are SaaS. Talk to us about deployment, or see pricing. The broader case for a cited data layer is in why AI agents need a trusted data layer.
Methodology
- Universe: every company that filed a form 10-K between January 1 and September 30 of 2023, 2024, 2025 and 2026, from the EDGAR quarterly form indexes, with SIC codes from EDGAR. One count per company per year. Banks are SIC 6000–6099; brokers and asset managers 6200–6299 (dealers, exchanges, investment advice) excluding 6221 commodity pools; insurers 6300–6499 excluding 6324 health insurers. SIC 61xx (credit institutions, asset-backed securities trusts and government-sponsored lenders) is excluded. The universe includes non-traded funds and BDCs that file 10-Ks.
- Mentions: EDGAR full-text search for "artificial intelligence" in any document of a 10-K or 10-K/A filed in the window. Generative AI means "generative AI" or "generative artificial intelligence". Full-text search returns every match for these queries; none hit its 10,000-result cap.
- Risk themes (2026 only): for the 462 companies with a mention, we parsed each company's latest original 10-K main document and located Item 1A as the longest span between an "Item 1A Risk Factors" heading and the next Item 1B, 1C or 2 heading (found in 405). We split the text into sentences and kept those with an AI term (artificial intelligence, AI, generative AI, machine learning, large language model, LLM). Each was tagged with keyword rules. Third-party AI vendors means vendor, service provider, supplier, outsourcing, or third-party AI, models, tools, software or services. Confidential-data exposure needs a sensitive-data term (confidential, proprietary or sensitive information, personal information, non-public information, client or customer data, trade secrets) and an exposure term (disclose, release, leak, expose, unauthorized, input, breach, compromise) in the same sentence. Shares in Figure 2 use the 384 companies with an AI sentence in a located Item 1A.
- Checks and limits: in hand-checked random samples, the vendor rule was right in 15 of 15 sentences and the data-exposure rule in 19 of 20. Keyword rules miss paraphrases, so shares are lower bounds, and the 57 filings where Item 1A could not be located are left out of the theme shares. A mention is a disclosure, not proof of adoption.
FAQ
What is an on-premise LLM?
An on-premise LLM is a language model that runs on hardware your firm owns and operates, so prompts and documents never leave your data center. It gives the most control and costs the most. It is also limited to models you can host yourself, such as open-weight models or the few proprietary models licensed for on-premise use; most frontier models are sold only as cloud endpoints.
What is the difference between private cloud, VPC and single-tenant deployment?
In a customer-VPC deployment, sometimes called private cloud, the vendor's software runs inside your own AWS, Azure or Google Cloud account, so you control the network and can see every outbound connection. In a single-tenant deployment, the vendor runs a dedicated instance for you in the vendor's cloud. Your data is isolated from other customers but still leaves your network. Either way, ask where the model itself runs.
Do regulators require banks or funds to run AI on-premise?
No. FINRA, the SEC, NYDFS, the Federal Reserve and the EU's DORA require firms to stay responsible for their vendors, know where data is stored and processed, and supervise their use of AI. NYDFS suggests contract clauses on the acceptable use of AI and on whether a firm's data may be used to train models. On-premise is one way to meet these expectations; a well-documented SaaS or VPC deployment is another.
How many financial firms mention AI in their 10-K?
In 10-Ks filed January to September 2026, 462 of the 550 banks, brokers, asset managers and insurers filing with the SEC (84%) mentioned artificial intelligence, up from 14% in 2023. 384 discussed AI in their risk factors, and 60% of those flagged reliance on third-party AI vendors, exposure of confidential data, or both.
Does Rogo offer on-premise or VPC deployment?
Rogo's website lists "Single Tenant Deployments" and says customer data is "stored in siloed environments, isolated from other customer data." On October 4, 2026 we found no public description of a customer-VPC or on-premise option. As with any vendor, ask what is available under contract.
What is FocusAlpha?
FocusAlpha is a SEC filings API and agent-ready financial data layer: it turns SEC filings (10-K, 10-Q, 8-K, 13F), earnings-call transcripts, and other trusted company communications into structured, normalized data where every value keeps its citation back to the source document. AI agents connect via API or MCP. On the Fund plan, FocusAlpha can be deployed inside your own VPC.