Wes Ellis./ a personal notebook
Technology. Stories. Side projects.
A few things worth writing down.
← Back to Script Library

SCRIPT LIBRARY · POWERSHELL

Count Manuscript Words by Chapter, Word Files Included

Get a word count for every chapter file and the whole manuscript, from Markdown, plain text or .docx, with a running total and a progress bar toward your target.

AT A GLANCEMeasure-ManuscriptWords.ps1
What it does
Counts the words in every .md, .txt and .docx chapter file in a folder, in natural chapter order, and returns one row per chapter with a running total, then a total row. Give it a target and you get the percentage and a progress bar too.
Requires
  • Windows PowerShell 5.1 or PowerShell 7+
  • No modules, and Word doesn't need to be installed
Permissions
Read access to the chapter files. Nothing is changed.
Runs on
Windows 10/11, macOS or Linux with PowerShell 7
Tested
Parse-checked and run in PowerShell 7.4 against Markdown with front matter, comments, links and a scene break, a plain-text file, a .docx with a tracked deletion, and a corrupt .docx, with and without -Target

Part 2 of the thread Tools for the writing desk

If you write one file per chapter, "how long is the book?" means opening every file, or pasting everything into one document just to look at the counter in the corner. And you want to know more than the total: which chapter is running long, which is thin, and how far you are from the number you promised yourself.

My first version of this was a quick loop that split on whitespace and printed a list. It was fine for Markdown and hopeless for Word files, which it read as raw bytes and counted as gibberish. This one reads .docx properly (it's a ZIP of XML inside) and cleans Markdown up before counting, so formatting marks, links and your notes-to-self in comments don't pad the number.

Output is objects, one per chapter, so you can sort them, chart them, or drop them into a spreadsheet to track progress week by week.

Measure-ManuscriptWords.ps1Download
<#
.SYNOPSIS
    Counts words per chapter file and for the whole manuscript, with an optional target.
.DESCRIPTION
    Reads every .md, .txt and .docx file in a folder, in natural order (Chapter 2 before Chapter 10), and returns
    one object per file with its word count and a running total, then a Total row.

    Markdown is cleaned up before counting: YAML front matter, HTML comments (handy for notes to yourself),
    link targets, image tags and formatting marks are ignored, so **bold** and [links](url) count as the words
    you'd read. Word documents are read straight from the .docx package (it's a ZIP of XML), so Word doesn't
    need to be installed; tracked deletions, comments and footnotes aren't counted. A "word" is any run of
    non-space characters with at least one letter or digit in it, so a lone dash or *** scene break doesn't count.

    Give it a -Target and each row also shows the percentage reached, and you get a progress bar at the end.
.PARAMETER Path
    A folder of chapter files, or specific files. Defaults to the current folder.
.PARAMETER Include
    Extensions to count. Default: .md, .txt, .docx.
.PARAMETER Recurse
    Include subfolders (for manuscripts split into part folders).
.PARAMETER Target
    Your target word count for the whole manuscript.
.PARAMETER ExcludeTotal
    Leave the Total row off, if you're feeding the output somewhere that sums it itself.
.EXAMPLE
    .\Measure-ManuscriptWords.ps1 -Path .\chapters -Target 80000
.EXAMPLE
    .\Measure-ManuscriptWords.ps1 -Path .\chapters | Export-Csv wordcount.csv -NoTypeInformation
#>
[CmdletBinding()]
param(
    [Parameter(ValueFromPipeline, ValueFromPipelineByPropertyName)]
    [Alias('FullName')]
    [string[]]$Path = '.',
    [string[]]$Include = @('.md', '.txt', '.docx'),
    [switch]$Recurse,
    [ValidateRange(1, 10000000)][int]$Target,
    [switch]$ExcludeTotal
)

begin {
    Add-Type -AssemblyName System.IO.Compression, System.IO.Compression.FileSystem
    $Include = @($Include | ForEach-Object { if ($_ -like '.*') { $_.ToLower() } else { ".$_".ToLower() } })
    $files = [System.Collections.Generic.List[IO.FileInfo]]::new()

    function Get-DocxText([string]$File) {
        $zip = [IO.Compression.ZipFile]::OpenRead($File)
        try {
            $entry = $zip.GetEntry('word/document.xml')
            if (-not $entry) { throw 'No word/document.xml inside. Is this really a .docx?' }
            $reader = [IO.StreamReader]::new($entry.Open())
            try { [xml]$xml = $reader.ReadToEnd() } finally { $reader.Dispose() }
        }
        finally { $zip.Dispose() }

        $ns = [Xml.XmlNamespaceManager]::new($xml.NameTable)
        $ns.AddNamespace('w', 'http://schemas.openxmlformats.org/wordprocessingml/2006/main')
        # One line per paragraph; w:t holds the visible text, w:tab and w:br separate words.
        $paragraphs = foreach ($p in $xml.SelectNodes('//w:body//w:p', $ns)) {
            ($p.SelectNodes('.//w:t | .//w:tab | .//w:br', $ns) | ForEach-Object { if ($_.LocalName -eq 't') { $_.InnerText } else { ' ' } }) -join ''
        }
        $paragraphs -join "`n"
    }

    function Get-MarkdownText([string]$Text) {
        $Text = $Text -replace '\A', ''
        $Text = $Text -replace '(?s)\A---\r?\n.*?\r?\n(---|\.\.\.)\r?\n', ''   # front matter
        $Text = $Text -replace '(?s)<!--.*?-->', ''                              # comments and notes to self
        $Text = $Text -replace '!\[[^\]]*\]\([^)]*\)', ''                         # images
        $Text = $Text -replace '\[([^\]]*)\]\([^)]*\)', '$1'                       # links: keep the text
        $Text = $Text -replace '<[^>]+>', ' '                                     # stray HTML tags
        $Text -replace '(?m)^\s{0,3}(#{1,6}|>|[-*+]|\d+\.)\s+', '' -replace '[*_~`]+', ''
    }

    function Measure-Word([string]$Text) {
        if (-not $Text) { return 0 }
        @($Text -split '\s+' | Where-Object { $_ -match '[\p{L}\p{N}]' }).Count
    }
}

process {
    foreach ($item in $Path) {
        if (Test-Path -LiteralPath $item -PathType Container) {
            Get-ChildItem -LiteralPath $item -File -Recurse:$Recurse | Where-Object { $Include -contains $_.Extension.ToLower() -and $_.Name -notlike '~$*' } | ForEach-Object { $files.Add($_) }
        }
        elseif (Test-Path -LiteralPath $item -PathType Leaf) { $files.Add((Get-Item -LiteralPath $item)) }
        else { Write-Warning "Not found: $item" }
    }
}

end {
    if (-not $files.Count) { Write-Warning 'No chapter files found.'; return }

    # Natural sort: pad every number so "Chapter 2" sorts before "Chapter 10".
    $sorted = $files | Sort-Object -Unique FullName | Sort-Object { $_.DirectoryName }, { [regex]::Replace($_.Name, '\d+', { $args[0].Value.PadLeft(10, '0') }) }
    $running = 0; $i = 0; $counted = 0

    foreach ($file in $sorted) {
        $i++
        Write-Progress -Activity 'Counting words' -Status $file.Name -PercentComplete (100 * $i / @($sorted).Count)
        $row = [ordered]@{ Chapter = $file.BaseName; File = $file.Name; Words = $null; RunningTotal = $null }
        try {
            $text = switch ($file.Extension.ToLower()) {
                '.docx' { Get-DocxText $file.FullName }
                '.md'   { Get-MarkdownText (Get-Content -LiteralPath $file.FullName -Raw -Encoding UTF8 -ErrorAction Stop) }
                default { Get-Content -LiteralPath $file.FullName -Raw -Encoding UTF8 -ErrorAction Stop }
            }
            $row.Words = Measure-Word $text
            $running += $row.Words
            $counted++
        }
        catch {
            Write-Warning "$($file.Name): $($_.Exception.Message)"
        }
        $row.RunningTotal = $running
        if ($Target) { $row.PercentOfTarget = [math]::Round(100 * $running / $Target, 1) }
        [pscustomobject]$row
    }
    Write-Progress -Activity 'Counting words' -Completed

    if (-not $ExcludeTotal) {
        $total = [ordered]@{ Chapter = 'Total'; File = "$counted file(s)"; Words = $running; RunningTotal = $running }
        if ($Target) { $total.PercentOfTarget = [math]::Round(100 * $running / $Target, 1) }
        [pscustomobject]$total
    }

    if ($Target) {
        $pct = [math]::Min(1.0, $running / $Target)
        $filled = [int][math]::Round(30 * $pct)
        $left = [math]::Max(0, $Target - $running)
        $bar = '[' + ('#' * $filled) + ('-' * (30 - $filled)) + ']'
        Write-Host ('{0} {1:N0} of {2:N0} words ({3:N0}%){4}' -f $bar, $running, $Target, (100 * $running / $Target), $(if ($left) { ", $($left.ToString('N0')) to go" } else { '. Done!' }))
    }
}

Parameters

ParameterTypeDefaultWhat it's for
-Pathstring[].A folder of chapter files, or specific files. Takes pipeline input.
-Includestring[].md, .txt, .docxWhich extensions to count.
-Recurseswitch—Include subfolders, for books split into part folders.
-Targetint—Target word count for the whole manuscript. Adds a PercentOfTarget column and a progress bar.
-ExcludeTotalswitch—Leave off the Total row, if whatever you're feeding it adds its own.

Run it

Chapter counts and a bar toward 80,000 words.

.\Measure-ManuscriptWords.ps1 -Path .\chapters -Target 80000

Longest chapters first.

.\Measure-ManuscriptWords.ps1 -Path .\chapters -ExcludeTotal | Sort-Object Words -Descending | Select-Object -First 5

Append today's total to a progress log.

.\Measure-ManuscriptWords.ps1 -Path .\chapters | Where-Object Chapter -eq Total | Select-Object @{n='Date';e={Get-Date -Format yyyy-MM-dd}}, Words | Export-Csv progress.csv -Append -NoTypeInformation

Just the Word files in a drafts folder.

.\Measure-ManuscriptWords.ps1 -Path .\drafts -Include .docx

What you'll see

Example outputvalues are illustrative
Chapter     File             Words RunningTotal PercentOfTarget
-------     ----             ----- ------------ ---------------
Chapter 01  Chapter 01.md     3412         3412            4.3
Chapter 02  Chapter 02.md     2988         6400            8.0
Chapter 03  Chapter 03.docx   4105        10505           13.1
...
Chapter 18  Chapter 18.md     3890        61220           76.5
Total       18 file(s)       61220        61220           76.5

[#######################-------] 61,220 of 80,000 words (77%), 18,780 to go

How it works

  1. Collect the files. It gathers every file with an included extension from the folders you give it, skipping Word's ~$ lock files, then sorts them with every number padded out, so the running total follows your chapter order.
  2. Get plain text. Markdown has its front matter, comments, images, link targets, heading marks and emphasis characters stripped. A .docx is opened as a ZIP, and the text is pulled from word/document.xml: the w:t runs, with tabs and line breaks treated as spaces. Deleted text lives in w:delText, so it's skipped automatically.
  3. Count words. The text is split on whitespace, and only pieces with at least one letter or digit count. That keeps * * * scene breaks and em dashes out of the total.
  4. Return rows. Each chapter comes back with its count and the running total, plus PercentOfTarget if you set -Target. A file that can't be read gets a warning and a blank count, and the rest carry on.
  5. Draw the bar. With a target, the last thing it prints is a 30-character bar with the words left to go.

Take it further

  • Build the book next. When the counts look right, Convert-MarkdownToKdpHtml turns the same folder of chapters into a single Kindle-ready file.
  • Chart the pace. Log the Total row daily with the third example, then open the CSV in a spreadsheet for a words-per-day line.
  • Catch lopsided chapters. Pipe the rows to Measure-Object Words -Average -Maximum -Minimum to see which chapters are way off the average.

Things that'll trip you up

  • It won't match Word's number exactly. Every tool counts a little differently. A hyphenated word like "well-known" counts once here, as it does in Word, but a dash standing on its own between spaces doesn't count at all, and tools disagree on that one. Expect to be within a percent or so of Word or Scrivener, which is plenty for tracking progress.
  • Comments in Markdown don't count. Anything inside <!-- --> is ignored, along with YAML front matter. That's deliberate, so notes to yourself don't inflate the count. If you keep scene notes in plain text instead, they'll be counted.
  • Tracked changes count as accepted. In a .docx, inserted text counts and deleted text doesn't, as if you'd accepted everything. Word comments and footnotes aren't counted at all.
  • Name chapters so they sort. The natural sort handles "Chapter 2" before "Chapter 10", but a prologue named "Prologue" sorts after "Chapter". Prefix it with 00 if the running total should start there.