ACRAdvanced Cipher Research
Case study

Reconstructing vehicle journeys from fragmented probe data.

A materials and logistics client needed to understand where heavy vehicles travel from and to across a region, to plan capacity and sites. The only data available at scale was commercial floating-car probe data, which is designed to prevent exactly that reconstruction.

Posterior density narrowing

The problem

Probe data is sold as anonymised sessions: short runs of GPS pings with an identifier that rotates after a few minutes, a few kilometres, or at every stop depending on the panel. A single journey from a quarry to a construction site is delivered as a handful of unrelated fragments. Across the five panels we evaluated, schemas, sampling rates, rotation rules and vehicle mixes all differed, and one vendor resold several upstream sources under one label.

The client question, origin and destination pairs for heavy vehicles, cannot be answered without re-linking those fragments. That is a statistical association problem under heavy uncertainty, at a scale of millions of sessions a day.

What we built

1
Panel characterisation. Before any modelling, we measured each data source: identifier rotation behaviour, gap-length distributions, sampling jitter, the share of multi-file sessions, and the information ceiling each panel imposed. This ruled out two sources before money was spent on them.
2
Map-matching and routing. Every fragment is matched to the road network with a self-hosted routing engine. Fragment ends become road-network positions with free-flow travel-time estimates between them.
3
Candidate generation. For each fragment end, a space-time cone yields the set of fragment starts that could plausibly be the same vehicle. At central-London density that is around 120 candidates per end, with the true continuation present 99.5 percent of the time.
4
Likelihood scoring. Each candidate pair is scored by a likelihood combining the observed-to-free-flow travel-time ratio, a road-class-by-hour congestion prior learned from the data itself, speed continuity at the join, and hard gates from the panel's own rotation rules.
5
Ranking and abstention. Candidates are ranked by posterior; the margin between first and second determines whether the link is accepted or left open. The abstention threshold is tuned to the client's cost of a false join, not to a round number.
6
Evaluation harness. Journeys with known truth are planted into real day-scale data and the whole pipeline is scored on forced first-pick accuracy, true-rank distribution and margin. Every experiment is frozen, versioned and reported in plain English.

What we found

Where it stands

The pipeline runs on day-scale data for the whole of Great Britain, with a London development lab used to measure performance under the hardest density conditions. Work continues on per-segment speed priors and endpoint enrichment. The methods transfer directly to any linkage problem where identifiers rotate: devices, accounts, vessels, or sensors.

Figures are from the development lab in 2026 and describe method performance, not a client's operational data. Client and data vendors are not named.

Contact

Bring us a problem that resists the obvious model.

We take on a small number of engagements at a time. Tell us what you are measuring, what you need to decide, and what the data looks like today. We reply within two working days.

Email
info@advcipher.com

Office
IFZA, Dubai, United Arab Emirates
Working with clients in the UK, Europe and the Gulf.

What to include
The decision you need to support, the data sources you have (or want), timescales, and any constraints on privacy or deployment.

Start a conversation