Skip to content

Reliability and arrival models

Before booking, Bahnsparer evaluates the connection and its historical arrival. During the journey, DB's current time remains visible; the app adds only the “9 out of 10 by” value.

Contents · 7 topics

Which model provides each value

Reliability

Probability that every included transfer works, no train is cancelled and arrival at the destination is less than six minutes late.

Historical arrival

Typical delay and later 90th-percentile boundary for a connection before current operating data is available.

Arrival during the journey

Time by which nine out of ten journeys with a similar DB forecast actually arrived. DB's current time remains the primary value.

Data and time windows

Zuverlässigkeit
Juli 2024 bis Mai 2026; Abschlusstest Juni 2026 und Juli 2026
Historische Ankunft
Training bis Mai 2026; Abschlusstest Juni 2026 und Juli 2026
Ankunft während der Fahrt
Training Januar bis Mai 2025; Kalibrierung Juni 2025; Abschlusstest September 2025 und November 2025

The long-term models read monthly files containing scheduled time, reported time, train number, station and cancellation. The live model reads several historical forecast states for the same arrival.

The December timetable change alters train numbers and connections, so that month is excluded. Training, calibration, selection and final testing use separate months. Source files, code and dependencies are pinned with SHA-256.

Features used

Transfer

Historical arrival of the incoming train, departure of the connecting train, cancellations, scheduled buffer, time of day and weekend.

Historical destination arrival

Punctuality, median, 90th percentile and cancellations at the destination, plus time of day, weekend and scheduled duration.

Current destination arrival

DB destination delay, 15 to 360 minutes until scheduled arrival, train run, destination stop, time of day, weekend and regional/long-distance service.

Final-test metrics

Whole-journey reliability

Reiseketten
3.851.111
Brier-Fehler
0,2180
Log-Loss
0,6251
ROC AUC
0,7043
Kalibrierungsfehler
1,82 Prozentpunkte

Historical arrival

Mittlere Abweichung
3,21443 statt 3,22495 Minuten
90-Prozent-Abdeckung
90,755 Prozent
Modellgröße
87.324 Byte

Arrival during the journey

Geprüfte Prognosestände
636.556
Mittlere Abweichung
5,10 statt 5,48 Minuten
Verbesserung
7,0 Prozent
90-Prozent-Abdeckung
89,72 Prozent
Modellgröße
86.861 Byte

Approved and rejected candidates

A candidate competes with the shipped model on the same test rows. Checks cover total error, calibration, individual time and connection segments, model size and agreement between Python and the app's TypeScript evaluation.

The current data run replaced neither reliability nor historical arrival: both candidates lost relevant tests. The live model was approved only from 15 to 360 minutes before scheduled arrival. After that, DB itself was more accurate.

On-device evaluation

The app contains the compact tree structures as JSON and evaluates them directly on the device. There is no model server and no additional request per connection.

The reliability model uses 369,567 bytes; the historical arrival model uses 87,324 bytes und das Live-Modell 86.861 Byte. New models arrive in an app update and do not change without one.

Factors not included

The models know neither today's engineering work nor severe weather or the cause of a disruption. Current notices and cancellations remain separate DB information.

The 90th-percentile value during a journey is not a second official forecast. Alerts, ticket restrictions, sharing and alternative searches use only DB's current time.

For an explanation without test metrics, read What Bahnsparer knows about your arrival.