Skip to content

About

Python-based collection and analysis pipeline for PTA temporary speed restriction data

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

PTA TSR Collector

A Python-based collection and analysis pipeline for extracting Temporary Speed Restriction (TSR) data from Public Transport Authority of Western Australia (PTA) Weekly Notices.

The project discovers PTA Weekly Notice PDFs, downloads them, extracts the Current Temporary Speed Restrictions / Current Speed Restrictions table from each notice, stores the results in a resumable SQLite database, assigns recurring restrictions to master TSR records, and exports analysis-ready CSV files.


Tip

Eventual project goal A Power BI view that includes a map overlay of Perth's rail network, to see repeat hotspots for speed restrictions and maintenance issues.


Warning

Project still under active development; data set and analysis not yet complete, comprehensive and validated Version 2.4.3 adds targeted historical row recovery on top of the v2.4.2 reliability work, including header/footer/contact fragment suppression, spaced-speed normalisation, whole-line historical parsing, cleaner manual review output, and the same hard quality gates before analytics export.

The collector can be treated as an authoritative pipeline for the PTA TSR dataset, but any new extraction run should still be reviewed with the generated data-quality and rejection diagnostics before relying on it for conclusions.

Any new full run should be validated with the post-full-run checklist before Power BI is refreshed.


Table of contents

But what is a Temporary Speed Restriction (TSR)?

The Australian Rail Industry Standards Organisation (ARISO) explains the purpose and principal of Temporary Speed Restrictions in its publication Australian Network Rules And Procedures (ANRP) 3025 - Temporary Speed Restrictions.

However, good luck getting this document if you don't work for a member railway of ARISO.

In short - Temporary speed restrictions are operationally important because they can indicate infrastructure condition, defect remediation timelines, operational constraints, renewal backlogs, or recurring maintenance pressure points across a rail network.

Taking the description from ARC Infrastructure's Network Safeworking Rules and Procedures - Temporary Speed Restructions (Rule 3025):

The object of a TSR is to reduce the speed of Rail Traffic to ensure safe passage over a Section of Track when the Track is not safe for Normal Speeds.

It continues:

A TSR may be applied due to:

  • Infrastructure conditions;
  • risks to workers; or
  • weather conditions.

Which might give you an insight into why the public who uses passenger rail might be interested in TSR's - as these can have some of the biggest impact on our commutes, by slowing down trains from their regular speeds.

And, if you're a really inquisitive type - and have read Page 197 of the PTA Safeworking Rules and Procedures; you'll see that their Temporary Speed Restrictions rules (Rule 3025) is the exact same text as Rule 3025 found in ARC Infrastructure's Safeworking Rules and Procedures.

So no need to be a paying member of ARISO to read what Rule 3025 is.

Why this project exists

tl,dr: PTA TSR data isn't published in a format enabling analysis. Hence the script in this project is designed to extract and unlock the data from the PDF's, to then:

  • answer questions about issues impacting railway performance,
  • give visability to the data,
  • create usable information for citizens to drive Governments and Ministers to keep on top of Transport agencies and railway maintenance, and
  • help everyone have faster and smoother journeys with less delays (yes, it means you'll have less excuses for not being at work on time - but swings and roundabouts).

The PTA publishes Weekly Notices containing safety and operational information. Embedded in those notices is a table of current speed restrictions.

Whilst each weekly notice is useful on its own, the real analytical value comes from collecting the table across many weeks and asking questions such as:

  • Which TSRs have persisted the longest?
  • Which corridors, line sections, or directions see repeated TSRs?
  • Which restriction reasons appear most often?
  • How often do restrictions become effectively permanent?
  • How long does it take for different classes of rail infrastructure restrictions to be removed?
  • Are some restriction types seasonal, recurring, or clustered?

This script turns the temporary speed restrictions table contained within the weekly PDF notices into a longitudinal dataset that can support answering those kinds of questions.

Although this repository is tailored to the PTA's public document structure, the approach may be useful to anyone interested in analysing rail maintenance communications, operational notices, infrastructure performance signals, or public transport network impacts in another jurisdiction.

What the collector does

At a high level, the collector:

  • Connects to the PTA Safety Resources / Weekly Notices document browser.
  • Uses the PTA site's DNN Document Viewer JSON API to recursively enumerate folders.
  • Finds Weekly Notice PDFs.
  • Registers discovered PDFs in a local SQLite database.
  • Downloads each PDF only once.
  • Searches each PDF for the speed restriction table.
  • Extracts TSR rows from the table.
  • Normalises key fields such as notice date, location, distance, speed and date imposed.
  • Assigns every row a sequential occurrence ID.
  • Groups repeated weekly appearances of the same TSR under a master TSR ID.
  • Exports CSV files for spreadsheet, database, or BI analysis.

What changed in v2.4.3

Version 2.4.3 is the current targeted cleanup release on top of the v2.4.2 reliability and workflow update.

Major changes:

  • Repair-first extraction: rows are repaired and normalised before being rejected.
  • Parser artefact filtering: blank rows and low-content fragments are counted separately instead of being treated as rejected TSRs.
  • Multiple extraction strategies: table-line, table-text and text-line fallback extraction are scored and selected per page.
  • Speed text normalisation: variants such as 80km’h, 80kmh, 80kph, 80 km/h and 80 km / h are normalised to 80km/h.
  • Spaced-speed token repair: extracted text such as 8 0 k m / h is normalised to 80km/h before validation.
  • Chainage parsing improvements: decimal ranges such as 2.465 to 2.780 are accepted even when the source omits km after each number.
  • Named-location acceptance: historical rows such as Fremantle - Robbs Jetty Section are accepted when speed and reason are otherwise valid.
  • Whole-line historical repair: older shifted rows that collapse reason, location, speed, date and cancellation into one line can be repaired and accepted automatically.
  • Header/footer/contact fragment suppression: non-reviewable extraction debris is skipped before rejection and kept out of manual review templates.
  • Repeated historical pattern recovery: obvious recurring 2019-2023 patterns are auto-repaired where enough information is present.
  • Line and direction inference: common Perth network locations are mapped to inferred line names, line keys and directions where possible.
  • Manual review workflow: reviewable unresolved rows are exported to a user-editable CSV with an associated .txt instruction file.
  • Cleaner manual review output: only rows that remain plausibly reviewable after v2.4.3 suppression and repair are exported for user review.
  • Review reprocessing: reviewed rows can be applied with apply-review and then included in exported analytics.
  • Hard quality gates: analytics export is blocked when the current dataset is clearly unsafe to use, such as when the latest processed notice has zero accepted rows.

What is unique about this project

This is not a generic PDF scraper. Several PTA-specific and domain-specific issues are handled deliberately.

PTA document browser support

The PTA Safety Resources page does not expose Weekly Notice PDFs as ordinary static links in the initial HTML. The visible folder browser is driven by a DNN / DotNetNuke Document Viewer module.

The collector therefore uses the underlying API endpoint:

https://www.pta.wa.gov.au/API/DocumentViewer/ContentService/GetFolderContent

The collector starts at the identified Weekly Notices folder ID:

5160

and recursively walks child folder records returned by the API.

Mixed historical folder structures

The PTA Weekly Notice archive is not completely uniform. Some years contain PDFs directly under the year folder. Other years are split into month folders.

The collector does not assume one fixed folder depth. It follows every folder returned by the API and extracts PDF records wherever they are found.

Mixed historical filename formats

The notices use changing date styles, including examples such as:

Weekly Notice No. 26 Week Commencing 28th June 2026.pdf
Weekly Notice No. 44 - Week Ending 10 Nov 2018.pdf
Weekly Notice No. 35 - Week Ending 8th September 2018.pdf
Weekly Notice No. 31 Week Ending 6 August2022.pdf
Weekly Notice No. 04 Week Ending- 01st Febuary 2019.pdf

The collector attempts to normalise these into:

YYYY-MM-DD

Mixed historical table layouts

The TSR table has changed over time. Older notices can use headings such as:

Current Speed Restrictions

with columns like:

Section of Railway
Distance at or between (km)
Maximum Speed
Reason for Restriction
Date Imposed

Later notices can include:

Current Temporary Speed Restrictions (TSR)

with columns like:

Location To and From (km)
STN No.
Maximum Speed
Date Imposed
Reason for Restriction
Date to be Cancelled

The collector uses flexible table detection and column mapping rather than assuming one exact table schema.

Location and distance splitting

Current PTA notices may combine location, line code/direction, and kilometres into one field, for example:

MTMDN 7.900km - 8.150km
Beyond Service 21.012km - 26.000km

The collector splits these into separate fields where possible:

location
line_section
distance_km

This makes it easier to analyse restrictions by line, direction, section, or location.

Repair-first extraction

Earlier versions were too quick to quarantine rows. Version 2.4.3 continues the repair-first approach introduced in v2.4.2 and extends it with targeted historical whole-line recovery before rejection.

The extraction pipeline is:

PDF page
  -> identify likely TSR section
  -> attempt multiple extraction strategies
  -> score extraction strategies
  -> skip blank parser artefacts
  -> normalise OCR/text issues
  -> recover speed, location, date, reason and cancellation fields
  -> infer line, direction, chainage and reason group
  -> accept valid rows
  -> quarantine only meaningful unresolved rows

Parser artefact handling

PDF table extraction can detect row boundaries without successfully extracting text from cells. That can create rows such as:

["", "", "", "", "", ""]

Version 2.4.3 treats these as parser artefacts. Parser artefacts are not accepted rows, not rejected TSR rows, and not manual review rows. The current script also suppresses obvious header/footer/contact fragments before they enter the manual review path. Parser artefact counts are tracked separately through artifact_row_count, tsr_artifact_row, pta_tsr_source_pdfs.csv, and pta_tsr_data_quality_summary.csv.

Line and direction inference

The collector attempts to infer useful line and direction fields from line codes, corridor names and location text.

Examples:

Leederville - Stirling ... Down Main -> Joondalup Line / Down
Joondalup - Edgewater ... Down Main  -> Joondalup Line / Down
Fremantle - Shenton Park ... Up Main -> Fremantle Line / Up
Glen Iris to Cockburn ... Up & Down  -> Mandurah Line / Bidirectional
Kwinana to Wellard ... Down Main     -> Mandurah Line / Down
Nowergup Yard                        -> Joondalup Line / Yard

Where inference is not possible, the collector uses review-visible values such as UNCLASSIFIED, Unclassified, or Unknown rather than leaving key analytical fields blank.

TSR master matching

Every appearance of a TSR in a weekly notice is stored as an occurrence. The collector also creates a master record for the underlying TSR.

The current master matching rule is:

location + distance_km + stn_no + date_imposed

Speed is intentionally not part of the master identity. If a restriction changes from 80 km/h to 60 km/h, that is treated as a modification of the same TSR, not a brand-new TSR.

Repository contents

Suggested repository structure:

pta-tsr-collector/
  README.md
  pta_tsr_collector.py
  .gitignore

Generated runtime data is intentionally kept out of source control:

pta_tsr_data/
  pta_tsr.sqlite3
  pdfs/
  exports/
  logs/
  diagnostics/

Requirements

  • Windows, macOS, or Linux
  • Python 3.11 or newer recommended
  • Internet access to the PTA website
  • Python packages installed by the script when --install-deps is used:
    • requests
    • beautifulsoup4
    • pdfplumber

The project has been developed and tested primarily from Windows PowerShell.

Windows quick start

1. Install Python

Download Python from:

https://www.python.org/downloads/windows/

During installation, tick:

Add python.exe to PATH

2. Confirm Python works

Open PowerShell and run:

py --version

If py is not available, try:

python --version

3. Create the project folder

mkdir C:\PythonScripts\pta_tsr_collector
cd C:\PythonScripts\pta_tsr_collector

Place the script here:

C:\PythonScripts\pta_tsr_collector\pta_tsr_collector.py

4. Discover PDFs

py .\pta_tsr_collector.py discover --folder-id 5160 --install-deps

This should populate the local database with discovered Weekly Notice PDFs.

5. Check status

py .\pta_tsr_collector.py status

You should see a non-zero PDF count.

6. Process a small test batch

py .\pta_tsr_collector.py process --limit 5

This is a safer first processing test before running across the full archive.

7. Process the full archive

py .\pta_tsr_collector.py run --retry-failed

Common commands

Initialise database only

py .\pta_tsr_collector.py init

Discover PDFs only

py .\pta_tsr_collector.py discover --folder-id 5160

Process pending PDFs only

py .\pta_tsr_collector.py process

Retry failed or missing PDFs

py .\pta_tsr_collector.py process --retry-failed

Limit processing to a small batch

py .\pta_tsr_collector.py process --limit 10

Full discovery, processing and export run

py .\pta_tsr_collector.py run --retry-failed

Export CSVs again

py .\pta_tsr_collector.py export

Show current status

py .\pta_tsr_collector.py status

Create diagnostics folder

py .\pta_tsr_collector.py diagnostics

Output files

After running, the collector creates:

pta_tsr_data/
  pta_tsr.sqlite3
  pdfs/
  exports/
    pta_tsr_occurrences.csv
    pta_tsr_masters.csv
    pta_tsr_source_pdfs.csv
  logs/
    pta_tsr_collector.log
  diagnostics/
    dnn_folder_<folder_id>.json

pta_tsr_occurrences.csv

One row per TSR appearance in a Weekly Notice.

Important fields include:

tsr_record_id
tsr_master_id
notice_date
location
line_section
distance_km
stn_no
max_speed
date_imposed
reason
date_cancelled
source_pdf
source_url
source_page
source_row_number

pta_tsr_masters.csv

One row per underlying TSR after grouping repeated appearances across notices.

Important fields include:

tsr_master_id
first_seen_notice_date
last_seen_notice_date
weeks_seen
latest_location
latest_line_section
latest_distance_km
latest_stn_no
latest_max_speed
date_imposed
latest_reason
latest_date_cancelled
master_fingerprint

pta_tsr_source_pdfs.csv

Audit file for source PDF discovery and processing state.

Important fields include:

source_pdf_id
notice_date
filename
url
status
tsr_table_count
tsr_row_count
last_error

Manual review workflow

When manual review is needed

Manual review is needed when the collector finds a row with enough content to possibly be a TSR, but not enough certainty to accept automatically.

Examples include:

  • missing, contradictory or shifted fields;
  • a reason field that looks like a shifted cancellation token;
  • a row that contains speed/date information but lacks enough location or reason context;
  • an extraction result that may be valid but needs user confirmation.

Manual review is not intended for blank parser artefacts, header/footer fragments, or known contact-number debris.

Files generated for manual review

Run:

py .\pta_tsr_collector.py export-review-template

The script creates:

pta_tsr_data\review\manual_rejection_review_template.csv
pta_tsr_data\review\manual_rejection_review_template.txt

The .txt instruction file explains how to edit the CSV and how to reprocess reviewed rows after user action.

How to edit the manual review CSV

For each row:

  • Read cell_preview and raw_row_json.
  • If the row is a valid TSR, set accept_row to 1 and fill in the corrected fields.
  • If the row is not a valid TSR and should be ignored in future, set accept_row to 0.
  • If unsure, leave accept_row blank and optionally add review_notes.

Required fields when accept_row=1:

corrected_location
corrected_max_speed
corrected_reason

Recommended fields when available:

corrected_distance_km
corrected_stn_no
corrected_date_imposed
corrected_date_cancelled
corrected_line_key
corrected_line_name
corrected_location_direction

How to reprocess reviewed rows

After editing and saving the CSV, run:

py .\pta_tsr_collector.py apply-review --review-file .\pta_tsr_data\review\manual_rejection_review_template.csv
py .\pta_tsr_collector.py export

Rows with accept_row=1 are inserted into tsr_occurrence and included in exported CSVs. Rows with accept_row=0 are marked as ignored and excluded from future review templates. Rows with blank accept_row remain unresolved.

What not to manually review

Do not manually review parser artefacts such as:

["", "", "", "", "", ""]

Do not manually review standalone fragments unless there is enough surrounding context to reconstruct a complete TSR row, such as:

Direction
Up Main
Down Main
40.874km to
40.747km Up

Version 2.4.3 is designed to exclude these from the manual review template and count them separately as parser/header/footer artefacts.

Quality gates and data assurance

Version 2.4.3 refuses to export analytics when the data appears unsafe.

Blocking conditions include:

  • no accepted TSR rows exist;
  • the latest processed notice has zero accepted TSR rows;
  • accepted rows are missing required normalised fields such as line, direction, affected area or reason group.

If export is refused, run:

py .\pta_tsr_collector.py status
py .\pta_tsr_collector.py rejection-diagnostics

Then review:

pta_tsr_data\logs\pta_tsr_collector.log
pta_tsr_data\exports\pta_tsr_source_pdfs.csv
pta_tsr_data\diagnostics\rejection_review_compact_YYYYMMDD_HHMMSS\01_rejection_summary.csv
pta_tsr_data\diagnostics\rejection_review_compact_YYYYMMDD_HHMMSS\02_rejection_samples.csv

A high parser artefact count is not automatically a data error, but it is a signal that a layout-specific extractor path may need further tuning.

Post-full-run validation checklist

Run this checklist after a full historical backfill or after a major extractor update before refreshing Power BI or committing generated analytics files.

1. Confirm overall collector status

py .\pta_tsr_collector.py status

Review these values:

processed
failed
missing
TSR masters
TSR occurrences
Manual review rows
Parser artefacts skipped
Latest processed notice
Latest accepted TSR rows

Interpretation:

  • TSR occurrences should increase as PDFs are processed. If processed PDFs increase but accepted occurrences stop increasing, the extractor may be failing a layout.
  • Latest accepted TSR rows should normally be greater than zero when the latest Weekly Notice contains current TSRs.
  • Manual review rows should be plausible and should not be dominated by blank or near-blank extraction debris.
  • Parser artefacts skipped can be non-zero. Artefacts are extraction debris, not rejected TSRs, but a sudden spike in recent notices should be investigated.

2. Check rejection categories

Use this command to confirm that rejected rows are genuine manual-review rows rather than parser artefacts:

py -c "import sqlite3; c=sqlite3.connect('pta_tsr_data/pta_tsr.sqlite3'); c.row_factory=sqlite3.Row; [print(f'{r[0]} | {r[1]}: {r[2]}') for r in c.execute('select rejection_category, reject_reason, count(*) from tsr_rejected_row group by rejection_category, reject_reason order by count(*) desc')]"

Expected pattern:

manual_review_required | missing_or_invalid_speed: <count>
manual_review_required | missing_location_or_speed: <count>
manual_review_required | invalid_reason_or_shifted_columns: <count>

Interpretation of rejection categories:

  • manual_review_required means the row has enough content to potentially be a TSR but was not safe enough to accept automatically.
  • parser_artifact should generally not appear in tsr_rejected_row. Parser artefacts should be counted separately via artifact_row_count and tsr_artifact_row.

Interpretation of common rejection reasons:

  • missing_or_invalid_speed means the row did not contain a speed that could be confidently normalised, or the speed appeared in an unresolved layout pattern.
  • missing_location_or_speed means a required location or speed field was still missing after repair attempts.
  • invalid_reason_or_shifted_columns means the reason field looked like a shifted value, such as a cancellation token, direction fragment, date fragment, STN value, or other non-reason text.

3. Check recent processed PDF extraction counts

Use a parameterised SQLite query to avoid PowerShell/Python/SQL quoting problems:

py -c "import sqlite3; c=sqlite3.connect('pta_tsr_data/pta_tsr.sqlite3'); c.row_factory=sqlite3.Row; [print(f'{r[0]}: accepted={r[1]}, rejected={r[2]}, artefacts={r[3]}') for r in c.execute('select notice_date, tsr_row_count, rejected_row_count, artifact_row_count from source_pdf where status=? order by notice_date desc limit 20', ('processed',))]"

Do not use this broken form:

where status=''processed''

Inside a Python single-quoted string, ''processed'' is interpreted as adjacent Python string literals and becomes status=processed, which SQLite treats as a column name rather than the text value 'processed'.

Interpretation:

  • Recent notices should usually have non-zero accepted counts if the source PDF contains current TSRs.
  • rejected should be reviewed if it spikes on a recent notice.
  • After the current fragment-suppression fixes, recent notices can legitimately show non-zero artefacts while still having rejected=0. Those artefacts are commonly split header, footer or partial location fragments that are intentionally excluded from manual review.
  • High recent artefacts counts should still be investigated if accepted rows collapse, if the counts jump sharply compared with nearby notices, or if the artefacts look like genuine TSR rows rather than extraction debris.

4. Check parser artefacts separately

py -c "import sqlite3; c=sqlite3.connect('pta_tsr_data/pta_tsr.sqlite3'); c.row_factory=sqlite3.Row; [print(f'{r[0]}: {r[1]}') for r in c.execute('select notice_date, artifact_row_count from source_pdf where artifact_row_count > 0 order by notice_date desc limit 20')]"

Interpretation:

  • Historical artefacts are expected in some older layouts.
  • Recent/current artefacts do not automatically mean the extractor is failing. A healthy recent notice can still have non-zero artefacts if split header/footer/location fragments were suppressed correctly and accepted TSR counts remain plausible.
  • Recent/current artefact spikes should be investigated before refreshing Power BI when they rise suddenly, coincide with reduced accepted rows, or appear to include genuine TSR content.
  • Artefacts are not manual review rows unless the extractor preserved enough content to classify the row as manual_review_required.

5. Generate compact rejection diagnostics

py .\pta_tsr_collector.py rejection-diagnostics

This creates a timestamped folder under:

pta_tsr_data\diagnostics\rejection_review_compact_YYYYMMDD_HHMMSS\

Review these files:

01_rejection_summary.csv
02_rejection_samples.csv
03_manual_review_template.csv
03_manual_review_template.txt
README.txt

Use the three CSV files for Copilot review if further extractor tuning is needed. The .txt file explains local user action for the review template.

6. Inspect key analytics outputs before Power BI refresh

After a successful export, inspect:

pta_tsr_data\analytics\pta_tsr_active_current.csv
pta_tsr_data\analytics\pta_tsr_active_by_line.csv
pta_tsr_data\analytics\pta_tsr_active_by_cause.csv
pta_tsr_data\analytics\pta_tsr_data_quality_summary.csv

Minimum checks:

  • pta_tsr_active_current.csv should have a plausible number of rows for the latest processed Weekly Notice.
  • line_key, line_name, location_direction, affected_area and reason_group should be populated.
  • pta_tsr_active_by_line.csv should not collapse into a single blank line row.
  • pta_tsr_active_by_cause.csv should not collapse into a single blank reason row.
  • pta_tsr_data_quality_summary.csv should show accepted rows, manual-review rows and parser artefacts separately.

7. Decide whether manual review is needed

If manual_review_required rows remain and the rows matter for analysis, export or use the generated manual review template:

py .\pta_tsr_collector.py export-review-template

Edit:

pta_tsr_data\review\manual_rejection_review_template.csv

Read the associated instruction file before editing:

pta_tsr_data\review\manual_rejection_review_template.txt

After editing the CSV, apply corrections and regenerate exports:

py .\pta_tsr_collector.py apply-review --review-file .\pta_tsr_data\review\manual_rejection_review_template.csv
py .\pta_tsr_collector.py export

8. Only refresh Power BI after validation passes

Refresh Power BI only after:

  • export succeeds without a quality-gate error;
  • latest accepted TSR rows are plausible;
  • active-current rows are populated with line and reason fields;
  • rejection and artefact counts have been reviewed;
  • manual corrections, if any, have been applied and exports regenerated.

Processing statuses

The source_pdf table and pta_tsr_source_pdfs.csv use statuses such as:

discovered
downloaded
processed
duplicate
failed
missing

Meaning:

  • discovered — the PDF was found in the PTA document browser but has not yet been processed.
  • downloaded — the PDF was downloaded but not fully processed.
  • processed — the PDF was downloaded, scanned and exported successfully.
  • duplicate — the PDF was successfully downloaded but was identified as a duplicate source copy of another notice for the same notice_date and file content, so it is retained for traceability but excluded from tsr_occurrence.
  • failed — the script encountered an error while processing the PDF.
  • missing — the PTA API listed the PDF, but the file URL could not be downloaded, usually because of a stale or malformed PTA record.

Known real-world data issues

This project exists because the source material is useful but not analysis-ready. Known issues include:

  • Historical notices have inconsistent folder structures.
  • Historical notices have inconsistent file naming.
  • Some filenames contain typographical errors such as Febuary.
  • Some filenames omit spaces, such as August2022.
  • Some API-listed documents may return 404 Not Found when downloaded.
  • Table structures changed between earlier and later years.
  • Some older tables do not include all newer fields, such as STN number or cancellation date.
  • PDF table extraction can be affected by page layout, merged cells, repeated headers or subtle PDF formatting changes.

The collector is designed to continue processing despite these issues and to record diagnostics for later review.

Source Coverage Snapshot

After source-gap recovery and duplicate-source exclusion on 2026-07-05, the live dataset ended in this source-coverage state:

processed: 378
missing: 10
duplicate: 1
failed: 0

Interpretation:

  • The 10 missing rows are current PTA-side 404 Not Found gaps that remained unresolved after retrying the available URL normalisations.
  • The 1 duplicate row is a successfully downloaded duplicate source copy of an already processed notice. It is retained in source_pdf for traceability but does not contribute rows to tsr_occurrence.
  • The previously recoverable stale-URL and malformed-name cases were processed successfully and are now included in the dataset exports.

For a concise project record of these outcomes, see:

pta_tsr_data\diagnostics\source_gap_report_20260705.md

Data model overview

The collector uses a local SQLite database with three main tables.

source_pdf

Tracks every Weekly Notice PDF discovered from the PTA site.

tsr_occurrence

Stores every TSR row extracted from every Weekly Notice.

This is the most detailed table and is the best starting point for longitudinal analysis.

tsr_master

Stores grouped TSR identities across weeks.

This table supports duration analysis, such as first seen, last seen and weeks seen.

Example analysis questions

Once data has been collected, possible questions include:

  • Which TSRs have existed for the most weeks?
  • Which locations have the highest number of TSR occurrences?
  • Which restriction reasons are most common?
  • How many TSRs are listed as indefinite, permanent, TBA, TBC or TBD?
  • Which restrictions change speed while remaining active?
  • Which restrictions disappeared from notices after a planned cancellation date?
  • Which lines or sections have recurring restrictions?
  • Are restrictions clustered around particular kilometre ranges?
  • How does the number of active restrictions change over time?

Using this project for other rail networks

This project is PTA-specific in its discovery layer but more general in its concept.

If another operator publishes weekly or periodic notices containing speed restriction tables, the same pattern may work:

  • Identify where documents are published.
  • Determine whether the document index is static HTML, a JavaScript app, an API, SharePoint, an S3 bucket, or another document system.
  • Build a reliable document discovery layer.
  • Download documents to a local cache.
  • Extract the relevant table.
  • Store raw occurrences separately from grouped master records.
  • Export repeatable CSVs for analysis.

The most important architectural lesson is to separate:

document discovery
PDF download/cache
PDF table extraction
normalisation
master matching
CSV/report export

That separation makes the project easier to adapt when a website changes or when historical document formats vary.

Responsible use and interpretation

This project extracts data from public documents for analysis. The dataset should still be interpreted carefully.

A TSR appearing for a long period does not, by itself, prove neglect, poor maintenance, or unsafe operation. Long-running restrictions can result from many causes, including planned staged works, funding cycles, engineering constraints, access windows, risk controls, asset renewals, environmental conditions, or operational decisions.

The dataset is best treated as a structured evidence base for further investigation, not as a final conclusion.

Recommended practice:

  • keep source URLs and source PDFs for traceability;
  • distinguish missing data from genuine absence of restrictions;
  • review outliers manually against the original PDF;
  • avoid over-interpreting one field without operational context;
  • document any assumptions used in master matching or duration calculations.

Troubleshooting

SyntaxWarning: invalid escape sequence

Older versions of the script contained Windows paths in normal Python string literals. Version 2.2 uses raw strings for these help/docstring sections.

Found/registered 0 PDF link(s)

This means generic HTML crawling did not find PDFs. The PTA page uses a DNN Document Viewer API, so use:

py .\pta_tsr_collector.py discover --folder-id 5160

404 Not Found while downloading a PDF

Some PTA API records may point to stale or malformed file URLs. Version 2.2 tries conservative alternate URLs and then marks the record as missing if the PDF still cannot be downloaded.

CSV files are empty

Check status:

py .\pta_tsr_collector.py status

If PDFs are still discovered, run:

py .\pta_tsr_collector.py process --limit 5

A PDF fails table extraction

Create diagnostics:

py .\pta_tsr_collector.py diagnostics

Then inspect:

pta_tsr_data\logs\pta_tsr_collector.log
pta_tsr_data\exports\pta_tsr_source_pdfs.csv
pta_tsr_data\diagnostics\

Scheduling weekly updates on Windows

Once the full historical backfill is complete, the same script can be scheduled.

Suggested Windows Task Scheduler configuration:

  • Program/script:
py
  • Arguments:
C:\PythonScripts\pta_tsr_collector\pta_tsr_collector.py run --retry-failed
  • Start in:
C:\PythonScripts\pta_tsr_collector

Schedule the task after the weekly notice is normally published.

Git ignore suggestions

Generated data should usually not be committed:

pta_tsr_data/
__pycache__/
*.pyc
.env
.venv/

If you want to publish a sample dataset, consider creating a separate curated samples/ folder with small, documented extracts rather than committing the full local database and downloaded PDFs.

Limitations

  • The collector depends on the current PTA DNN Document Viewer API behaviour.
  • Historical PDFs may contain formatting changes not yet handled by the parser.
  • Automated table extraction may require further tuning after reviewing all historical notices.
  • Master matching is heuristic and should be reviewed before using outputs for formal conclusions.
  • The project does not yet include a Power BI dashboard or interactive reporting layer.

Roadmap ideas

Potential future improvements:

  • stronger validation of extracted rows;
  • improved parsing of line codes and directions;
  • explicit active/inactive TSR lifecycle table;
  • automatic weekly change reports;
  • Power BI or DuckDB-ready model;
  • anomaly detection for unusually long-running TSRs;
  • HTML or Markdown summary reports;
  • unit tests with sampled historical PDFs;
  • GitHub Actions linting and smoke tests;
  • configuration file support for adapting the collector to other networks.

Licence and attribution

This project is an independent data extraction and analysis tool. It is not affiliated with, endorsed by, or maintained by the Public Transport Authority of Western Australia.

Users are responsible for complying with applicable website terms, copyright rules, public-sector information reuse requirements, and any local policies that apply to their use case.

Acknowledgement

This project was created to make public operational notice data easier to analyse over time. The aim is to support transparent, evidence-based discussion about rail infrastructure condition, maintenance impacts, and operational restrictions.

About

Python-based collection and analysis pipeline for PTA temporary speed restriction data

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages