Original research · Easy Sentence
The Plain English Paradox
Since 1998 the SEC has required plain English in part of every prospectus — the summary and the risk factors — and nowhere else in it. Unlike the Plain Writing Act, this rule has teeth. We measured 984 prospectuses on both sides of that internal boundary. The sections the rule covers are the least plain part of the document.
- Prospectuses
- 984
- Words analysed
- 17.2M
- Period
- 1996–2001
- Result
- Covered sections score worse
A boundary inside a single document
Our first study asked whether the Plain Writing Act of 2010 changed how federal agencies write, and found no detectable effect. That Act has no enforcement mechanism. This rule does: the SEC issues comment letters when a filing fails it, and a registration statement does not go effective while a comment letter is outstanding.
17 CFR 230.421(d) · adopted 63 FR 6370 · effective 1 October 1998
"To enhance the readability of the prospectus, you must use plain English principles in the organization, language, and design of the front and back cover pages, the summary, and the risk factors section."
The rule stops there. Business, Management's Discussion and Analysis, Use of Proceeds, Management — the whole back half of the prospectus — carry no such obligation. So the treatment and the control are not different documents by different authors in different years. They are different pages of the same document, written by the same people, for the same offering, on the same day, with the same lawyers.
Summary · Risk factors
Named in 421(d). Must "substantially comply" with six plain English principles, and the SEC checks.
Business · MD&A
Same document, same authors, same day. No plain English obligation at all.
That design is the point. A difference-in-differences has to assume the two arms would have moved in parallel; the first study's did not, which is why its headline was a chart rather than a coefficient. Here the primary estimate needs no such assumption. Company, industry, era, counsel and house style are held constant by construction.
The regulated sections are the hardest ones
On almost every structural measure the rule names, the covered sections score worse than the uncovered ones in the same document. Sentences in the summary and risk factors run 2.0 words longer on average than sentences in Business and MD&A — and the interval is nowhere near zero.
| Measure | 421(d) principle | Gap | 95% interval | Excludes zero? |
|---|---|---|---|---|
| Average sentence length | Short sentences | 1.98 | [1.73, 2.26] | yes |
| Sentences over the long-sentence threshold | Short sentences | 0.021 | [0.016, 0.026] | yes |
| Sentences containing a passive | Active voice | 0.003 | [-0.001, 0.008] | no |
| Stacked negatives | No multiple negatives | 0.65 | [0.61, 0.69] | yes |
| Nominalizations | Everyday words | 1.45 | [0.92, 1.97] | yes |
| Wordy phrases | Everyday words / no jargon | 1.59 | [1.51, 1.67] | yes |
This is not the result a reader expects, so it is worth being precise about what it does and does not say. It does not say the rule made the writing worse. It says that after more than twenty years of an enforced plain English mandate, the only sections legally required to be plain remain the densest prose in the document.
There is an obvious candidate explanation, and it is not the rule: risk factors are intrinsically about contingency and legal exposure, and that is hard to write simply. A Business section describes what a company does. A risk factors section describes what might go wrong, conditionally, without conceding anything. Genre could account for the level.
Which is exactly why the level is not the interesting number. The interesting number is whether the gap moved when the rule took effect — because genre did not change on 1 October 1998, and the rule did.
The gap does narrow — and we cannot give the rule the credit
Before the rule, covered sections ran 2.8 words per sentence longer than uncovered ones. After it, 1.0. The gap closed by 1.7 words.
Measured as a difference-in-differences, the change is -1.74 words [-2.41, -1.20]. Read on its own, that is the result the study was hoping for: a plain-language mandate with enforcement behind it, doing what the unenforced one did not.
We are not reporting it that way, for the same reason the first study did not report its apparently-significant result. A difference-in-differences only carries a causal reading if the arms were moving in parallel beforehand.
What the pre-trend does to the story
Fitted across the 9 quarters before the rule was published, the gap was already moving at 0.69 words per year. Projected forward over the same span the difference-in-differences covers, that pre-existing trend accounts for 79% of the estimated change.
5 of the 6 headline measures fail that test — average sentence length, sentences over the long-sentence threshold, stacked negatives, nominalizations, wordy phrases. For those, a straight continuation of what was already happening explains most of what happens at the boundary, and this design cannot separate the two.
One measure passes: sentences containing a passive, where the pre-period was flat and the change at the boundary is -0.018 [-0.029, -0.009]. Flesch Reading Ease passes too, and more strongly: its pre-trend runs against the post-rule move (-0.31 points per year before, a 4.5-point improvement after), so the trend cannot be what produced it.
We report those as suggestive rather than conclusive. Two clean event studies inside a set of confounded ones is a weak basis for a causal claim, and singling them out after the fact is exactly how a study talks itself into a result. Note also that the measure with the cleanest event study — passive voice — is the one measure whose within-document gap is not distinguishable from zero.
So the honest reading is narrower than the headline would like. The within-document gap is real, large, and robust. Its narrowing across 1998 is real. Attributing that narrowing to Rule 421(d) is supported on two measures out of seven and confounded on the rest — prospectus writing was already changing, and the rule arrived while it changed.
One further caveat that cuts against us, and belongs here rather than in a footnote: the study window is short. The rule was adopted in February 1998 and bit that October, so the pre-period is barely two years. That is why the pre-trend is fitted on quarters rather than years, and it is why a longer window would be the single most useful extension of this work.
Where the estimate does hold up
The within-document gap survives every reasonable change to how it is measured:
| Specification | Filings | Gap | 95% interval |
|---|---|---|---|
| Primary (business + MD&A as control) | 984 | 1.98 | [1.73, 2.26] |
| All other uncovered sections as control | 984 | -3.91 | [-4.30, -3.53] |
| Business + MD&A only, dropping "The Company" | 928 | 1.90 | [1.71, 2.10] |
| Filings with under 25% absorbed sub-headings | 619 | 2.04 | [1.79, 2.30] |
| Plain-text filings only | 961 | 1.98 | [1.72, 2.28] |
| 424B4 filings only | 525 | 2.23 | [1.84, 2.69] |
One row flips sign, and it is the informative one. Measured against all remaining uncovered sections — underwriting, tax, legal matters, experts — the covered sections come out ahead. Those sections are pure boilerplate and are denser than anything else in the document. Risk factors sit between narrative business prose and legal boilerplate: worse than the writing, better than the fine print.
Whose risk factors are hardest to read
Separately from the question above — and making no claim about the rule — here is how each industry writes the sections the SEC regulates. This is a description of the corpus, not an estimate of anything.
HOTELS & MOTELS filings average 30.9 words per sentence in their covered sections. SEMICONDUCTORS & RELATED DEVICES filings: 23.6.
| Industry (SIC) | Filings | Avg. sentence | Gap vs uncovered | Flesch |
|---|---|---|---|---|
| HOTELS & MOTELS 7011 | 11 | 30.9 | 2.9 | 24.6 |
| SAVINGS INSTITUTION, FEDERALLY CHARTERED 6035 | 12 | 28.7 | 2.0 | 27.3 |
| BLANK CHECKS 6770 | 12 | 28.7 | 1.1 | 29.4 |
| REAL ESTATE INVESTMENT TRUSTS 6798 | 12 | 28.4 | 1.5 | 28.6 |
| MORTGAGE BANKERS & LOAN CORRESPONDENTS 6162 | 8 | 27.8 | 2.2 | 32.0 |
| CABLE & OTHER PAY TELEVISION SERVICES 4841 | 11 | 27.7 | 1.8 | 27.8 |
| SURGICAL & MEDICAL INSTRUMENTS & APPARATUS 3841 | 11 | 27.5 | 3.1 | 23.8 |
| RETAIL-EATING PLACES 5812 | 15 | 27.5 | 2.6 | 27.1 |
| TELEPHONE COMMUNICATIONS (NO RADIO TELEPHONE) 4813 | 26 | 26.8 | 0.8 | 27.5 |
| PHARMACEUTICAL PREPARATIONS 2834 | 21 | 26.3 | 2.2 | 24.9 |
| CRUDE PETROLEUM & NATURAL GAS 1311 | 14 | 26.3 | 1.7 | 30.4 |
| TELEPHONE & TELEGRAPH APPARATUS 3661 | 12 | 26.2 | 2.4 | 26.3 |
| NATIONAL COMMERCIAL BANKS 6021 | 13 | 26.2 | 1.9 | 26.6 |
| STATE COMMERCIAL BANKS 6022 | 24 | 26.2 | 1.6 | 30.1 |
| RETAIL-CATALOG & MAIL-ORDER HOUSES 5961 | 12 | 26.0 | 1.7 | 27.2 |
| SERVICES-COMPUTER PROGRAMMING SERVICES 7371 | 31 | 25.8 | 2.1 | 27.4 |
| SERVICES-COMPUTER PROCESSING & DATA PREPARATION 7374 | 15 | 25.6 | 1.8 | 29.1 |
| SERVICES-COMPUTER PROGRAMMING, DATA PROCESSING, ETC. 7370 | 8 | 25.4 | 1.6 | 30.9 |
| SERVICES-COMPUTER INTEGRATED SYSTEMS DESIGN 7373 | 26 | 25.3 | 1.1 | 28.4 |
| WHOLESALE-COMPUTER & PERIPHERAL EQUIPMENT & SOFTWARE 5045 | 9 | 25.3 | 0.8 | 25.7 |
| MISCELLANEOUS ELECTRICAL MACHINERY, EQUIPMENT & SUPPLIES 3690 | 8 | 25.3 | 1.5 | 31.1 |
| SERVICES-PREPACKAGED SOFTWARE 7372 | 57 | 25.3 | 1.9 | 27.7 |
| SERVICES-MANAGEMENT CONSULTING SERVICES 8742 | 13 | 25.2 | 1.0 | 30.0 |
| COMMUNICATION SERVICES, NEC 4899 | 8 | 25.2 | 0.4 | 31.7 |
| SERVICES-ADVERTISING 7310 | 9 | 25.0 | 0.6 | 27.8 |
| SERVICES-COMMERCIAL PHYSICAL & BIOLOGICAL RESEARCH 8731 | 19 | 24.3 | 1.2 | 28.4 |
| SERVICES-BUSINESS SERVICES, NEC 7389 | 53 | 24.2 | 1.3 | 33.3 |
| RADIO & TV BROADCASTING & COMMUNICATIONS EQUIPMENT 3663 | 8 | 23.9 | 2.0 | 28.1 |
| BIOLOGICAL PRODUCTS (NO DIAGNOSTIC SUBSTANCES) 2836 | 14 | 23.7 | 0.9 | 27.2 |
| SEMICONDUCTORS & RELATED DEVICES 3674 | 29 | 23.6 | 1.5 | 29.7 |
The eight-filing floor is doing real work and is disclosed for that reason. Without it the table is topped by whichever industry happened to file twice, and a two-filing industry says nothing about how an industry writes.
How this was measured
The population is every 424B1 and 424B4 filing — final offering prospectuses — in EDGAR's quarterly indexes for 1996 through 2001: 6,861 filings after folding co-registrant duplicates. From that, 1,440 were sampled by a deterministic hash of the accession number, stratified by quarter, so the sample reproduces exactly rather than being random. SEC filings are public records.
Each filing is split into sections by heading, and each section is classified as covered by 421(d) or not. That split is the entire study, and it is the part most likely to be wrong, so every filing's arm word counts are published alongside its scores.
Every measure comes from an analyzer that already runs on this site — the same code behind the passive voice finder, the sentence length visualizer, and the plain language checker. One rule was written for this study, because Rule 421(d)(2)(vi) names "no multiple negatives" as a principle and nothing on the site measured it; it ships as an ordinary product rule, not as study-only code.
Confidence intervals come from a filer-clustered bootstrap (1,000 replicates, seeded so they reproduce exactly). Several prospectuses from one company share a house style and a law firm, and treating them as independent observations manufactures precision that is not there.
What this cannot tell you
- Cover pages are excluded from the covered arm, though the rule covers them. They are mandated legends and a price table rather than authored prose, and measuring them would score boilerplate the drafter did not choose.
- One of the six principles is not measured at all. "Tabular presentation or bullet lists for complex material" is a layout instruction, not a property of prose, and this study says so rather than substituting a proxy for it.
- Not every prospectus qualifies. A filing joins the comparison only if both arms are present and reach 200 words. Many 424B filings legitimately have no Business or MD&A section: a secondary offering by an already-public company incorporates them by reference from its 10-K. That is a property of the document, not an extraction failure, but it means the sample is offerings that restate their own business.
- A higher share of filings qualifies before the rule than after it — 78% against 62%. This is the study's most serious weakness, and it is the same failure mode that killed the first study's original plan: if the mix of documents shifts at exactly the boundary under test, a composition change can be mistaken for an effect. The qualifying filings do look alike on either side — 389 filings from 351 filers before against 483 from 445 after, similar arm sizes, similar form-type mix — but "the survivors look alike" is weaker evidence than "nothing was lost", and the difference is one more reason the before/after estimate is reported as secondary.
- Readability formulas are crude, which is why the headline measures here are structural facts rather than composite grades. See how readability formulas fail.
Check our work
Everything below is the actual output of the analysis, not a summary of it. The per-filing file is the one to start with: it carries both arms' scores side by side, so the headline number can be recomputed from two columns.
- Per-filing scores — 984 rows, both arms, every measure, plus the accession number to fetch the original from EDGAR
- Event study — the covered/uncovered gap by year, every metric
- Industry ranking — the table above, in full
Published August 27, 2026. Filings are public records; the analysis is free to reuse with attribution. Corrections are welcome — get in touch.