A Python-based collection and analysis pipeline for extracting Temporary Speed Restriction (TSR) data from Public Transport Authority of Western Australia (PTA) Weekly Notices.
The project discovers PTA Weekly Notice PDFs, downloads them, extracts the Current Temporary Speed Restrictions / Current Speed Restrictions table from each notice, stores the results in a resumable SQLite database, assigns recurring restrictions to master TSR records, and exports analysis-ready CSV files.
Tip
Eventual project goal A Power BI view that includes a map overlay of Perth's rail network, to see repeat hotspots for speed restrictions and maintenance issues.
Warning
Project still under active development; data set and analysis not yet complete, comprehensive and validated Version 2.4.3 adds targeted historical row recovery on top of the v2.4.2 reliability work, including header/footer/contact fragment suppression, spaced-speed normalisation, whole-line historical parsing, cleaner manual review output, and the same hard quality gates before analytics export.
The collector can be treated as an authoritative pipeline for the PTA TSR dataset, but any new extraction run should still be reviewed with the generated data-quality and rejection diagnostics before relying on it for conclusions.
Any new full run should be validated with the post-full-run checklist before Power BI is refreshed.
- But what is a Temporary Speed Restriction (TSR)?
- Why this project exists
- What the collector does
- What changed in v2.4.3
- What is unique about this project
- Repository contents
- Requirements
- Windows quick start
- Common commands
- Output files
- Manual review workflow
- Quality gates and data assurance
- Post-full-run validation checklist
- 1. Confirm overall collector status
- 2. Check rejection categories
- 3. Check recent processed PDF extraction counts
- 4. Check parser artefacts separately
- 5. Generate compact rejection diagnostics
- 6. Inspect key analytics outputs before Power BI refresh
- 7. Decide whether manual review is needed
- 8. Only refresh Power BI after validation passes
- Processing statuses
- Known real-world data issues
- Data model overview
- Example analysis questions
- Using this project for other rail networks
- Responsible use and interpretation
- Troubleshooting
- Scheduling weekly updates on Windows
- Git ignore suggestions
- Limitations
- Roadmap ideas
- Licence and attribution
- Acknowledgement
The Australian Rail Industry Standards Organisation (ARISO) explains the purpose and principal of Temporary Speed Restrictions in its publication Australian Network Rules And Procedures (ANRP) 3025 - Temporary Speed Restrictions.
However, good luck getting this document if you don't work for a member railway of ARISO.
In short - Temporary speed restrictions are operationally important because they can indicate infrastructure condition, defect remediation timelines, operational constraints, renewal backlogs, or recurring maintenance pressure points across a rail network.
Taking the description from ARC Infrastructure's Network Safeworking Rules and Procedures - Temporary Speed Restructions (Rule 3025):
The object of a TSR is to reduce the speed of Rail Traffic to ensure safe passage over a Section of Track when the Track is not safe for Normal Speeds.
It continues:
A TSR may be applied due to:
- Infrastructure conditions;
- risks to workers; or
- weather conditions.
Which might give you an insight into why the public who uses passenger rail might be interested in TSR's - as these can have some of the biggest impact on our commutes, by slowing down trains from their regular speeds.
And, if you're a really inquisitive type - and have read Page 197 of the PTA Safeworking Rules and Procedures; you'll see that their Temporary Speed Restrictions rules (Rule 3025) is the exact same text as Rule 3025 found in ARC Infrastructure's Safeworking Rules and Procedures.
So no need to be a paying member of ARISO to read what Rule 3025 is.
tl,dr: PTA TSR data isn't published in a format enabling analysis. Hence the script in this project is designed to extract and unlock the data from the PDF's, to then:
- answer questions about issues impacting railway performance,
- give visability to the data,
- create usable information for citizens to drive Governments and Ministers to keep on top of Transport agencies and railway maintenance, and
- help everyone have faster and smoother journeys with less delays (yes, it means you'll have less excuses for not being at work on time - but swings and roundabouts).
The PTA publishes Weekly Notices containing safety and operational information. Embedded in those notices is a table of current speed restrictions.
Whilst each weekly notice is useful on its own, the real analytical value comes from collecting the table across many weeks and asking questions such as:
- Which TSRs have persisted the longest?
- Which corridors, line sections, or directions see repeated TSRs?
- Which restriction reasons appear most often?
- How often do restrictions become effectively permanent?
- How long does it take for different classes of rail infrastructure restrictions to be removed?
- Are some restriction types seasonal, recurring, or clustered?
This script turns the temporary speed restrictions table contained within the weekly PDF notices into a longitudinal dataset that can support answering those kinds of questions.
Although this repository is tailored to the PTA's public document structure, the approach may be useful to anyone interested in analysing rail maintenance communications, operational notices, infrastructure performance signals, or public transport network impacts in another jurisdiction.
At a high level, the collector:
- Connects to the PTA Safety Resources / Weekly Notices document browser.
- Uses the PTA site's DNN Document Viewer JSON API to recursively enumerate folders.
- Finds Weekly Notice PDFs.
- Registers discovered PDFs in a local SQLite database.
- Downloads each PDF only once.
- Searches each PDF for the speed restriction table.
- Extracts TSR rows from the table.
- Normalises key fields such as notice date, location, distance, speed and date imposed.
- Assigns every row a sequential occurrence ID.
- Groups repeated weekly appearances of the same TSR under a master TSR ID.
- Exports CSV files for spreadsheet, database, or BI analysis.
Version 2.4.3 is the current targeted cleanup release on top of the v2.4.2 reliability and workflow update.
Major changes:
- Repair-first extraction: rows are repaired and normalised before being rejected.
- Parser artefact filtering: blank rows and low-content fragments are counted separately instead of being treated as rejected TSRs.
- Multiple extraction strategies: table-line, table-text and text-line fallback extraction are scored and selected per page.
- Speed text normalisation: variants such as
80km’h,80kmh,80kph,80 km/hand80 km / hare normalised to80km/h. - Spaced-speed token repair: extracted text such as
8 0 k m / his normalised to80km/hbefore validation. - Chainage parsing improvements: decimal ranges such as
2.465 to 2.780are accepted even when the source omitskmafter each number. - Named-location acceptance: historical rows such as
Fremantle - Robbs Jetty Sectionare accepted when speed and reason are otherwise valid. - Whole-line historical repair: older shifted rows that collapse reason, location, speed, date and cancellation into one line can be repaired and accepted automatically.
- Header/footer/contact fragment suppression: non-reviewable extraction debris is skipped before rejection and kept out of manual review templates.
- Repeated historical pattern recovery: obvious recurring 2019-2023 patterns are auto-repaired where enough information is present.
- Line and direction inference: common Perth network locations are mapped to inferred line names, line keys and directions where possible.
- Manual review workflow: reviewable unresolved rows are exported to a user-editable CSV with an associated
.txtinstruction file. - Cleaner manual review output: only rows that remain plausibly reviewable after v2.4.3 suppression and repair are exported for user review.
- Review reprocessing: reviewed rows can be applied with
apply-reviewand then included in exported analytics. - Hard quality gates: analytics export is blocked when the current dataset is clearly unsafe to use, such as when the latest processed notice has zero accepted rows.
This is not a generic PDF scraper. Several PTA-specific and domain-specific issues are handled deliberately.
The PTA Safety Resources page does not expose Weekly Notice PDFs as ordinary static links in the initial HTML. The visible folder browser is driven by a DNN / DotNetNuke Document Viewer module.
The collector therefore uses the underlying API endpoint:
https://www.pta.wa.gov.au/API/DocumentViewer/ContentService/GetFolderContent
The collector starts at the identified Weekly Notices folder ID:
5160
and recursively walks child folder records returned by the API.
The PTA Weekly Notice archive is not completely uniform. Some years contain PDFs directly under the year folder. Other years are split into month folders.
The collector does not assume one fixed folder depth. It follows every folder returned by the API and extracts PDF records wherever they are found.
The notices use changing date styles, including examples such as:
Weekly Notice No. 26 Week Commencing 28th June 2026.pdf
Weekly Notice No. 44 - Week Ending 10 Nov 2018.pdf
Weekly Notice No. 35 - Week Ending 8th September 2018.pdf
Weekly Notice No. 31 Week Ending 6 August2022.pdf
Weekly Notice No. 04 Week Ending- 01st Febuary 2019.pdf
The collector attempts to normalise these into:
YYYY-MM-DD
The TSR table has changed over time. Older notices can use headings such as:
Current Speed Restrictions
with columns like:
Section of Railway
Distance at or between (km)
Maximum Speed
Reason for Restriction
Date Imposed
Later notices can include:
Current Temporary Speed Restrictions (TSR)
with columns like:
Location To and From (km)
STN No.
Maximum Speed
Date Imposed
Reason for Restriction
Date to be Cancelled
The collector uses flexible table detection and column mapping rather than assuming one exact table schema.
Current PTA notices may combine location, line code/direction, and kilometres into one field, for example:
MTMDN 7.900km - 8.150km
Beyond Service 21.012km - 26.000km
The collector splits these into separate fields where possible:
location
line_section
distance_km
This makes it easier to analyse restrictions by line, direction, section, or location.
Earlier versions were too quick to quarantine rows. Version 2.4.3 continues the repair-first approach introduced in v2.4.2 and extends it with targeted historical whole-line recovery before rejection.
The extraction pipeline is:
PDF page
-> identify likely TSR section
-> attempt multiple extraction strategies
-> score extraction strategies
-> skip blank parser artefacts
-> normalise OCR/text issues
-> recover speed, location, date, reason and cancellation fields
-> infer line, direction, chainage and reason group
-> accept valid rows
-> quarantine only meaningful unresolved rows
PDF table extraction can detect row boundaries without successfully extracting text from cells. That can create rows such as:
["", "", "", "", "", ""]Version 2.4.3 treats these as parser artefacts. Parser artefacts are not accepted rows, not rejected TSR rows, and not manual review rows. The current script also suppresses obvious header/footer/contact fragments before they enter the manual review path. Parser artefact counts are tracked separately through artifact_row_count, tsr_artifact_row, pta_tsr_source_pdfs.csv, and pta_tsr_data_quality_summary.csv.
The collector attempts to infer useful line and direction fields from line codes, corridor names and location text.
Examples:
Leederville - Stirling ... Down Main -> Joondalup Line / Down
Joondalup - Edgewater ... Down Main -> Joondalup Line / Down
Fremantle - Shenton Park ... Up Main -> Fremantle Line / Up
Glen Iris to Cockburn ... Up & Down -> Mandurah Line / Bidirectional
Kwinana to Wellard ... Down Main -> Mandurah Line / Down
Nowergup Yard -> Joondalup Line / Yard
Where inference is not possible, the collector uses review-visible values such as UNCLASSIFIED, Unclassified, or Unknown rather than leaving key analytical fields blank.
Every appearance of a TSR in a weekly notice is stored as an occurrence. The collector also creates a master record for the underlying TSR.
The current master matching rule is:
location + distance_km + stn_no + date_imposed
Speed is intentionally not part of the master identity. If a restriction changes from 80 km/h to 60 km/h, that is treated as a modification of the same TSR, not a brand-new TSR.
Suggested repository structure:
pta-tsr-collector/
README.md
pta_tsr_collector.py
.gitignore
Generated runtime data is intentionally kept out of source control:
pta_tsr_data/
pta_tsr.sqlite3
pdfs/
exports/
logs/
diagnostics/
- Windows, macOS, or Linux
- Python 3.11 or newer recommended
- Internet access to the PTA website
- Python packages installed by the script when
--install-depsis used:requestsbeautifulsoup4pdfplumber
The project has been developed and tested primarily from Windows PowerShell.
Download Python from:
https://www.python.org/downloads/windows/
During installation, tick:
Add python.exe to PATH
Open PowerShell and run:
py --versionIf py is not available, try:
python --versionmkdir C:\PythonScripts\pta_tsr_collector
cd C:\PythonScripts\pta_tsr_collectorPlace the script here:
C:\PythonScripts\pta_tsr_collector\pta_tsr_collector.py
py .\pta_tsr_collector.py discover --folder-id 5160 --install-depsThis should populate the local database with discovered Weekly Notice PDFs.
py .\pta_tsr_collector.py statusYou should see a non-zero PDF count.
py .\pta_tsr_collector.py process --limit 5This is a safer first processing test before running across the full archive.
py .\pta_tsr_collector.py run --retry-failedpy .\pta_tsr_collector.py initpy .\pta_tsr_collector.py discover --folder-id 5160py .\pta_tsr_collector.py processpy .\pta_tsr_collector.py process --retry-failedpy .\pta_tsr_collector.py process --limit 10py .\pta_tsr_collector.py run --retry-failedpy .\pta_tsr_collector.py exportpy .\pta_tsr_collector.py statuspy .\pta_tsr_collector.py diagnosticsAfter running, the collector creates:
pta_tsr_data/
pta_tsr.sqlite3
pdfs/
exports/
pta_tsr_occurrences.csv
pta_tsr_masters.csv
pta_tsr_source_pdfs.csv
logs/
pta_tsr_collector.log
diagnostics/
dnn_folder_<folder_id>.json
One row per TSR appearance in a Weekly Notice.
Important fields include:
tsr_record_id
tsr_master_id
notice_date
location
line_section
distance_km
stn_no
max_speed
date_imposed
reason
date_cancelled
source_pdf
source_url
source_page
source_row_number
One row per underlying TSR after grouping repeated appearances across notices.
Important fields include:
tsr_master_id
first_seen_notice_date
last_seen_notice_date
weeks_seen
latest_location
latest_line_section
latest_distance_km
latest_stn_no
latest_max_speed
date_imposed
latest_reason
latest_date_cancelled
master_fingerprint
Audit file for source PDF discovery and processing state.
Important fields include:
source_pdf_id
notice_date
filename
url
status
tsr_table_count
tsr_row_count
last_error
Manual review is needed when the collector finds a row with enough content to possibly be a TSR, but not enough certainty to accept automatically.
Examples include:
- missing, contradictory or shifted fields;
- a reason field that looks like a shifted cancellation token;
- a row that contains speed/date information but lacks enough location or reason context;
- an extraction result that may be valid but needs user confirmation.
Manual review is not intended for blank parser artefacts, header/footer fragments, or known contact-number debris.
Run:
py .\pta_tsr_collector.py export-review-templateThe script creates:
pta_tsr_data\review\manual_rejection_review_template.csv
pta_tsr_data\review\manual_rejection_review_template.txt
The .txt instruction file explains how to edit the CSV and how to reprocess reviewed rows after user action.
For each row:
- Read
cell_previewandraw_row_json. - If the row is a valid TSR, set
accept_rowto1and fill in the corrected fields. - If the row is not a valid TSR and should be ignored in future, set
accept_rowto0. - If unsure, leave
accept_rowblank and optionally addreview_notes.
Required fields when accept_row=1:
corrected_location
corrected_max_speed
corrected_reason
Recommended fields when available:
corrected_distance_km
corrected_stn_no
corrected_date_imposed
corrected_date_cancelled
corrected_line_key
corrected_line_name
corrected_location_direction
After editing and saving the CSV, run:
py .\pta_tsr_collector.py apply-review --review-file .\pta_tsr_data\review\manual_rejection_review_template.csv
py .\pta_tsr_collector.py exportRows with accept_row=1 are inserted into tsr_occurrence and included in exported CSVs. Rows with accept_row=0 are marked as ignored and excluded from future review templates. Rows with blank accept_row remain unresolved.
Do not manually review parser artefacts such as:
["", "", "", "", "", ""]Do not manually review standalone fragments unless there is enough surrounding context to reconstruct a complete TSR row, such as:
Direction
Up Main
Down Main
40.874km to
40.747km Up
Version 2.4.3 is designed to exclude these from the manual review template and count them separately as parser/header/footer artefacts.
Version 2.4.3 refuses to export analytics when the data appears unsafe.
Blocking conditions include:
- no accepted TSR rows exist;
- the latest processed notice has zero accepted TSR rows;
- accepted rows are missing required normalised fields such as line, direction, affected area or reason group.
If export is refused, run:
py .\pta_tsr_collector.py status
py .\pta_tsr_collector.py rejection-diagnosticsThen review:
pta_tsr_data\logs\pta_tsr_collector.log
pta_tsr_data\exports\pta_tsr_source_pdfs.csv
pta_tsr_data\diagnostics\rejection_review_compact_YYYYMMDD_HHMMSS\01_rejection_summary.csv
pta_tsr_data\diagnostics\rejection_review_compact_YYYYMMDD_HHMMSS\02_rejection_samples.csv
A high parser artefact count is not automatically a data error, but it is a signal that a layout-specific extractor path may need further tuning.
Run this checklist after a full historical backfill or after a major extractor update before refreshing Power BI or committing generated analytics files.
py .\pta_tsr_collector.py statusReview these values:
processed
failed
missing
TSR masters
TSR occurrences
Manual review rows
Parser artefacts skipped
Latest processed notice
Latest accepted TSR rows
Interpretation:
TSR occurrencesshould increase as PDFs are processed. If processed PDFs increase but accepted occurrences stop increasing, the extractor may be failing a layout.Latest accepted TSR rowsshould normally be greater than zero when the latest Weekly Notice contains current TSRs.Manual review rowsshould be plausible and should not be dominated by blank or near-blank extraction debris.Parser artefacts skippedcan be non-zero. Artefacts are extraction debris, not rejected TSRs, but a sudden spike in recent notices should be investigated.
Use this command to confirm that rejected rows are genuine manual-review rows rather than parser artefacts:
py -c "import sqlite3; c=sqlite3.connect('pta_tsr_data/pta_tsr.sqlite3'); c.row_factory=sqlite3.Row; [print(f'{r[0]} | {r[1]}: {r[2]}') for r in c.execute('select rejection_category, reject_reason, count(*) from tsr_rejected_row group by rejection_category, reject_reason order by count(*) desc')]"Expected pattern:
manual_review_required | missing_or_invalid_speed: <count>
manual_review_required | missing_location_or_speed: <count>
manual_review_required | invalid_reason_or_shifted_columns: <count>
Interpretation of rejection categories:
manual_review_requiredmeans the row has enough content to potentially be a TSR but was not safe enough to accept automatically.parser_artifactshould generally not appear intsr_rejected_row. Parser artefacts should be counted separately viaartifact_row_countandtsr_artifact_row.
Interpretation of common rejection reasons:
missing_or_invalid_speedmeans the row did not contain a speed that could be confidently normalised, or the speed appeared in an unresolved layout pattern.missing_location_or_speedmeans a required location or speed field was still missing after repair attempts.invalid_reason_or_shifted_columnsmeans the reason field looked like a shifted value, such as a cancellation token, direction fragment, date fragment, STN value, or other non-reason text.
Use a parameterised SQLite query to avoid PowerShell/Python/SQL quoting problems:
py -c "import sqlite3; c=sqlite3.connect('pta_tsr_data/pta_tsr.sqlite3'); c.row_factory=sqlite3.Row; [print(f'{r[0]}: accepted={r[1]}, rejected={r[2]}, artefacts={r[3]}') for r in c.execute('select notice_date, tsr_row_count, rejected_row_count, artifact_row_count from source_pdf where status=? order by notice_date desc limit 20', ('processed',))]"Do not use this broken form:
where status=''processed''Inside a Python single-quoted string, ''processed'' is interpreted as adjacent Python string literals and becomes status=processed, which SQLite treats as a column name rather than the text value 'processed'.
Interpretation:
- Recent notices should usually have non-zero
acceptedcounts if the source PDF contains current TSRs. rejectedshould be reviewed if it spikes on a recent notice.- After the current fragment-suppression fixes, recent notices can legitimately show non-zero
artefactswhile still havingrejected=0. Those artefacts are commonly split header, footer or partial location fragments that are intentionally excluded from manual review. - High recent
artefactscounts should still be investigated if accepted rows collapse, if the counts jump sharply compared with nearby notices, or if the artefacts look like genuine TSR rows rather than extraction debris.
py -c "import sqlite3; c=sqlite3.connect('pta_tsr_data/pta_tsr.sqlite3'); c.row_factory=sqlite3.Row; [print(f'{r[0]}: {r[1]}') for r in c.execute('select notice_date, artifact_row_count from source_pdf where artifact_row_count > 0 order by notice_date desc limit 20')]"Interpretation:
- Historical artefacts are expected in some older layouts.
- Recent/current artefacts do not automatically mean the extractor is failing. A healthy recent notice can still have non-zero artefacts if split header/footer/location fragments were suppressed correctly and accepted TSR counts remain plausible.
- Recent/current artefact spikes should be investigated before refreshing Power BI when they rise suddenly, coincide with reduced accepted rows, or appear to include genuine TSR content.
- Artefacts are not manual review rows unless the extractor preserved enough content to classify the row as
manual_review_required.
py .\pta_tsr_collector.py rejection-diagnosticsThis creates a timestamped folder under:
pta_tsr_data\diagnostics\rejection_review_compact_YYYYMMDD_HHMMSS\
Review these files:
01_rejection_summary.csv
02_rejection_samples.csv
03_manual_review_template.csv
03_manual_review_template.txt
README.txt
Use the three CSV files for Copilot review if further extractor tuning is needed. The .txt file explains local user action for the review template.
After a successful export, inspect:
pta_tsr_data\analytics\pta_tsr_active_current.csv
pta_tsr_data\analytics\pta_tsr_active_by_line.csv
pta_tsr_data\analytics\pta_tsr_active_by_cause.csv
pta_tsr_data\analytics\pta_tsr_data_quality_summary.csv
Minimum checks:
pta_tsr_active_current.csvshould have a plausible number of rows for the latest processed Weekly Notice.line_key,line_name,location_direction,affected_areaandreason_groupshould be populated.pta_tsr_active_by_line.csvshould not collapse into a single blank line row.pta_tsr_active_by_cause.csvshould not collapse into a single blank reason row.pta_tsr_data_quality_summary.csvshould show accepted rows, manual-review rows and parser artefacts separately.
If manual_review_required rows remain and the rows matter for analysis, export or use the generated manual review template:
py .\pta_tsr_collector.py export-review-templateEdit:
pta_tsr_data\review\manual_rejection_review_template.csv
Read the associated instruction file before editing:
pta_tsr_data\review\manual_rejection_review_template.txt
After editing the CSV, apply corrections and regenerate exports:
py .\pta_tsr_collector.py apply-review --review-file .\pta_tsr_data\review\manual_rejection_review_template.csv
py .\pta_tsr_collector.py exportRefresh Power BI only after:
exportsucceeds without a quality-gate error;- latest accepted TSR rows are plausible;
- active-current rows are populated with line and reason fields;
- rejection and artefact counts have been reviewed;
- manual corrections, if any, have been applied and exports regenerated.
The source_pdf table and pta_tsr_source_pdfs.csv use statuses such as:
discovered
downloaded
processed
duplicate
failed
missing
Meaning:
discovered— the PDF was found in the PTA document browser but has not yet been processed.downloaded— the PDF was downloaded but not fully processed.processed— the PDF was downloaded, scanned and exported successfully.duplicate— the PDF was successfully downloaded but was identified as a duplicate source copy of another notice for the samenotice_dateand file content, so it is retained for traceability but excluded fromtsr_occurrence.failed— the script encountered an error while processing the PDF.missing— the PTA API listed the PDF, but the file URL could not be downloaded, usually because of a stale or malformed PTA record.
This project exists because the source material is useful but not analysis-ready. Known issues include:
- Historical notices have inconsistent folder structures.
- Historical notices have inconsistent file naming.
- Some filenames contain typographical errors such as
Febuary. - Some filenames omit spaces, such as
August2022. - Some API-listed documents may return
404 Not Foundwhen downloaded. - Table structures changed between earlier and later years.
- Some older tables do not include all newer fields, such as STN number or cancellation date.
- PDF table extraction can be affected by page layout, merged cells, repeated headers or subtle PDF formatting changes.
The collector is designed to continue processing despite these issues and to record diagnostics for later review.
After source-gap recovery and duplicate-source exclusion on 2026-07-05, the live dataset ended in this source-coverage state:
processed: 378
missing: 10
duplicate: 1
failed: 0
Interpretation:
- The
10 missingrows are current PTA-side404 Not Foundgaps that remained unresolved after retrying the available URL normalisations. - The
1 duplicaterow is a successfully downloaded duplicate source copy of an already processed notice. It is retained insource_pdffor traceability but does not contribute rows totsr_occurrence. - The previously recoverable stale-URL and malformed-name cases were processed successfully and are now included in the dataset exports.
For a concise project record of these outcomes, see:
pta_tsr_data\diagnostics\source_gap_report_20260705.md
The collector uses a local SQLite database with three main tables.
Tracks every Weekly Notice PDF discovered from the PTA site.
Stores every TSR row extracted from every Weekly Notice.
This is the most detailed table and is the best starting point for longitudinal analysis.
Stores grouped TSR identities across weeks.
This table supports duration analysis, such as first seen, last seen and weeks seen.
Once data has been collected, possible questions include:
- Which TSRs have existed for the most weeks?
- Which locations have the highest number of TSR occurrences?
- Which restriction reasons are most common?
- How many TSRs are listed as indefinite, permanent, TBA, TBC or TBD?
- Which restrictions change speed while remaining active?
- Which restrictions disappeared from notices after a planned cancellation date?
- Which lines or sections have recurring restrictions?
- Are restrictions clustered around particular kilometre ranges?
- How does the number of active restrictions change over time?
This project is PTA-specific in its discovery layer but more general in its concept.
If another operator publishes weekly or periodic notices containing speed restriction tables, the same pattern may work:
- Identify where documents are published.
- Determine whether the document index is static HTML, a JavaScript app, an API, SharePoint, an S3 bucket, or another document system.
- Build a reliable document discovery layer.
- Download documents to a local cache.
- Extract the relevant table.
- Store raw occurrences separately from grouped master records.
- Export repeatable CSVs for analysis.
The most important architectural lesson is to separate:
document discovery
PDF download/cache
PDF table extraction
normalisation
master matching
CSV/report export
That separation makes the project easier to adapt when a website changes or when historical document formats vary.
This project extracts data from public documents for analysis. The dataset should still be interpreted carefully.
A TSR appearing for a long period does not, by itself, prove neglect, poor maintenance, or unsafe operation. Long-running restrictions can result from many causes, including planned staged works, funding cycles, engineering constraints, access windows, risk controls, asset renewals, environmental conditions, or operational decisions.
The dataset is best treated as a structured evidence base for further investigation, not as a final conclusion.
Recommended practice:
- keep source URLs and source PDFs for traceability;
- distinguish missing data from genuine absence of restrictions;
- review outliers manually against the original PDF;
- avoid over-interpreting one field without operational context;
- document any assumptions used in master matching or duration calculations.
Older versions of the script contained Windows paths in normal Python string literals. Version 2.2 uses raw strings for these help/docstring sections.
This means generic HTML crawling did not find PDFs. The PTA page uses a DNN Document Viewer API, so use:
py .\pta_tsr_collector.py discover --folder-id 5160Some PTA API records may point to stale or malformed file URLs. Version 2.2 tries conservative alternate URLs and then marks the record as missing if the PDF still cannot be downloaded.
Check status:
py .\pta_tsr_collector.py statusIf PDFs are still discovered, run:
py .\pta_tsr_collector.py process --limit 5Create diagnostics:
py .\pta_tsr_collector.py diagnosticsThen inspect:
pta_tsr_data\logs\pta_tsr_collector.log
pta_tsr_data\exports\pta_tsr_source_pdfs.csv
pta_tsr_data\diagnostics\
Once the full historical backfill is complete, the same script can be scheduled.
Suggested Windows Task Scheduler configuration:
- Program/script:
py
- Arguments:
C:\PythonScripts\pta_tsr_collector\pta_tsr_collector.py run --retry-failed
- Start in:
C:\PythonScripts\pta_tsr_collector
Schedule the task after the weekly notice is normally published.
Generated data should usually not be committed:
pta_tsr_data/
__pycache__/
*.pyc
.env
.venv/If you want to publish a sample dataset, consider creating a separate curated samples/ folder with small, documented extracts rather than committing the full local database and downloaded PDFs.
- The collector depends on the current PTA DNN Document Viewer API behaviour.
- Historical PDFs may contain formatting changes not yet handled by the parser.
- Automated table extraction may require further tuning after reviewing all historical notices.
- Master matching is heuristic and should be reviewed before using outputs for formal conclusions.
- The project does not yet include a Power BI dashboard or interactive reporting layer.
Potential future improvements:
- stronger validation of extracted rows;
- improved parsing of line codes and directions;
- explicit active/inactive TSR lifecycle table;
- automatic weekly change reports;
- Power BI or DuckDB-ready model;
- anomaly detection for unusually long-running TSRs;
- HTML or Markdown summary reports;
- unit tests with sampled historical PDFs;
- GitHub Actions linting and smoke tests;
- configuration file support for adapting the collector to other networks.
This project is an independent data extraction and analysis tool. It is not affiliated with, endorsed by, or maintained by the Public Transport Authority of Western Australia.
Users are responsible for complying with applicable website terms, copyright rules, public-sector information reuse requirements, and any local policies that apply to their use case.
This project was created to make public operational notice data easier to analyse over time. The aim is to support transparent, evidence-based discussion about rail infrastructure condition, maintenance impacts, and operational restrictions.