Why Asia Has No EDGAR: EDINET, DART and MOPS

Ask a US-centric research stack a question about a Taiwanese supplier and you usually get silence. The instinct is to blame vendor laziness, but the real reason is structural: there is no Asian EDGAR to plug into. Japan splits its disclosure across two systems, Korea consolidates into one, Taiwan runs its own portal, and the official text arrives in three different languages under different accounting taxonomies. That's an integration problem, not a scaling problem — and it's worth solving, because the return evidence for cross-market disclosure signals is now measurable.

TLDR:

  • No single equivalent of EDGAR: Japan splits EDINET (statutory, FSA) from TDnet (timely, TSE); Korea's DART is a single submission point; Taiwan publishes via MOPS.
  • Language is a data problem: Taiwan's official announcement text is Traditional Chinese only; TDnet defaults to Japanese; DART is predominantly Korean — though Korea is expanding English disclosure.
  • The payoff is documented: in NUS's CrossAlpha benchmark, cross-market disclosure-derived peers beat domestic baselines in US→Japan by ICIR 0.39 vs 0.07–0.18.

Why there is no Asian EDGAR: four markets, four systems

The systems aren't variations on a theme. They're organised on different principles, which is the first thing that breaks a "just point it at the other market" plan.

Market System Operator Notes
US EDGAR SEC Single federal filing system
Japan EDINET FSA Statutory filings; XBRL with J-GAAP taxonomy since 2008
Japan TDnet Tokyo Stock Exchange Separate timely disclosure channel
Korea DART Financial Supervisory Service One filing serves FSC, FSS and KRX
Taiwan MOPS TWSE Statutory portal; machine access via TWSE OpenAPI

Sources: EDINET API materials (FSA); DART structure; MOPS / TWSE OpenAPI; TDnet.

Comparison of disclosure structures: the US has one system (EDGAR); Japan splits into EDINET for statutory filings and TDnet for timely disclosure; Korea consolidates into DART as a single submission point serving FSC, FSS and KRX; Taiwan publishes through MOPS with machine access via TWSE OpenAPI

Figure 1: There is no Asian EDGAR.

Two consequences are easy to miss. In Japan, a company's material news and its statutory filings live in different places, so covering EDINET alone leaves you blind to timely disclosure. In Korea the opposite holds — DART's single-submission design means one integration reaches filings that serve three regulators at once, which makes it comparatively tractable.

Language is a data problem, not a UX problem

This is the part teams underestimate. It isn't that the interface is in another language; it's that the legally operative text is.

Market Language of record English availability
Taiwan Traditional Chinese only (official text under the disclosure mandate) Limited
Japan Japanese by default on TDnet Only a small subset of large caps post bilingual titles
Korea Predominantly Korean; detailed and timely filings often Korean-first Expanding — an open platform now offers 83 disclosure data types in English

Sources: MOPS/TWSE; TDnet; Rutgers research guide; XBRL.org on Korea's English expansion.

Korea is the encouraging case: it has expanded its English disclosure system explicitly to attract foreign investment, with an open platform now offering 83 disclosure data types in English. But "English is available" and "English is the record" are different claims, and for research you have to know which one you're relying on.

The hard part isn't fetching — it's comparability

Even with every file in hand, the data isn't yet usable side by side, and the best articulation of why comes from the academic work. In CrossAlpha, a National University of Singapore benchmark for cross-market factor research, the authors state the obstacle plainly: building such a benchmark is hard because "filings differ across languages and regulatory systems." They break it into three layers.

Layer The obstacle
Disclosure Severe filing heterogeneity across languages and regulatory systems
Firm links Raw text similarity is biased by shared reporting style and industry jargon
Evaluation Mismatched time zones and trading calendars introduce look-ahead bias unless execution lags are handled

Source: CrossAlpha (arXiv 2605.29286), NUS — abstract and introduction.

Their remedy for the first layer is instructive: an LLM-based "Disclosure Distillation" stage that converts diverse annual reports into a shared ten-category English business schema. In other words, comparability had to be manufactured — you cannot compare a J-GAAP XBRL filing to a Traditional Chinese announcement without first mapping both into a common representation.

Pipeline: filings arrive in Japanese XBRL under J-GAAP, Korean under K-GAAP or IFRS, and Traditional Chinese announcements; a distillation step maps them into one shared English schema; the result is comparable firms, with each value still linked back to its original filing

Figure 2: Comparability has to be manufactured — and the original document stays the citation.

Why it's worth the trouble

The benchmark also answers the "so what." CrossAlpha spans about 3,600 firms and 10,700 firm-year reports across the US, Japan, Taiwan, South Korea and Hong Kong, paired with 11 years of daily prices. Its headline result is that links derived from disclosure text predict cross-market returns better than the usual proxies.

Peer-construction method (US→Japan) ICIR
Cross-market disclosure-derived peers 0.39
Domestic text, industry-code, return-correlation baselines 0.07–0.18

Source: CrossAlpha — up to a five-fold improvement in predictive skill. The paper also reports that filtering graph-retrieved neighbours with an LLM agent nearly doubled portfolio Sharpe in event-driven spillover trading.

That reframes APAC disclosure coverage. It isn't a completeness checkbox for firms that happen to look at Asia — the evidence suggests the cross-border links themselves carry signal that domestic data can't reproduce. A supply-chain relationship disclosed in a Taiwanese annual report is information about a US semiconductor name, and vice versa.

What a data layer has to do

Three requirements fall out of the above, and they're the same ones we'd hold any source to. Normalize into a common schema so the same field means the same thing in every market. Keep the original filing as the citation, so a normalized value can always be traced back to the document it came from — translation and mapping must be auditable, not a black box. And respect calendars and time zones, because cross-market timing errors quietly become look-ahead bias.

That's the standard behind the US disclosure coverage we run today, and it's the one we carried into Asia coverage — Japan, Korea, Taiwan, and China and Hong Kong. If your research stops at the US border, it's worth knowing the border is an engineering artefact, not a fact about where the information is.

FAQ

Is there an Asian equivalent of SEC EDGAR?

Not a single one. Japan splits statutory filings (EDINET, run by the FSA) from timely disclosure (TDnet, run by the Tokyo Stock Exchange); Korea consolidates into DART under the Financial Supervisory Service; Taiwan uses MOPS with machine access via the TWSE OpenAPI. Covering "Asia" means integrating several systems.

What language are Asian corporate filings published in?

The language of record differs by market. Taiwan's official announcement text is Traditional Chinese; Japan's TDnet defaults to Japanese with only a small subset of large caps posting bilingual titles; Korea's DART is predominantly Korean, although Korea has expanded English disclosure with an open platform offering 83 disclosure data types in English.

Why is cross-market disclosure data hard to use even after you download it?

Because comparability has to be built. As the NUS CrossAlpha benchmark puts it, filings "differ across languages and regulatory systems," textual similarity is biased by shared reporting style, and mismatched time zones and trading calendars introduce look-ahead bias unless execution lags are handled explicitly.

Do cross-market disclosure signals actually predict returns?

The evidence says yes. In CrossAlpha's US→Japan setting, peers derived from disclosure text achieved an ICIR of 0.39 versus 0.07–0.18 for domestic text, industry-code, and return-correlation baselines — up to a five-fold improvement in predictive skill across a benchmark of ~3,600 firms in the US, Japan, Taiwan, South Korea and Hong Kong.

What is FocusAlpha?

FocusAlpha is a SEC filings API and agent-ready financial data layer: it turns SEC filings (10-K, 10-Q, 8-K, 13F), earnings-call transcripts, and other trusted company communications into structured, normalized data where every value keeps its citation back to the source document. Coverage extends beyond the US to Japan, Korea, Taiwan, China and Hong Kong. AI agents connect via API or MCP to research public companies from complete, trusted information.

FOR AGENTS

This post is available as plain markdown with structured metadata — no scraping required.

GET .md →