ACRAdvanced Cipher Research
Statistical research

Models built to be validated, not just fitted.

We work in the space between textbook statistics and messy operational data: identifiers that rotate, sensors that drop out, populations that are skewed and heavy-tailed, and decisions that cannot wait for perfect information.

Stacked density curves being refined over time

Bayesian inference

Bayesian methods let us state what we knew before the data arrived, update it coherently, and report what remains uncertain. We use them wherever a decision depends on the uncertainty as much as the estimate.

  • Hierarchical models that share strength across groups, regions or time periods without pretending they are identical.
  • Sequential updating for streams: posteriors carried forward as new observations arrive, with explicit forgetting where the world drifts.
  • Posterior computation by Markov chain Monte Carlo where fidelity matters, and by variational approximation where throughput matters, with diagnostics for both.
  • Priors derived from physical constraints and historical data, documented so a reviewer can disagree with them.

In practice: a travel-time prior for a road link, learned from millions of observed traversals, gives a likelihood for whether two fragments belong to the same vehicle. The posterior over candidate matches is what we rank, and the posterior mass left on the second candidate is what tells us when to abstain.

Multivariate and skew-normal likelihoods

Real populations are rarely symmetric. Speeds, dwell times, incomes, claim sizes and demand all have long tails and skew. Forcing them into Gaussian assumptions produces confident, wrong answers.

  • Multivariate skew-normal and skew-t families for population modelling, with likelihood optimisation that stays numerically stable in high dimensions.
  • Covariance and precision matrix estimation with shrinkage and sparsity where the number of variables rivals the number of observations.
  • Mixture models for populations with distinct sub-groups, with identifiability checked rather than assumed.
  • Copula constructions when marginals are well understood but the dependence structure is the question.

In practice: the covariance matrices our population models produce are not a by-product. They are the deliverable: they tell a client which measured quantities move together and which only appear to.

Stochastic processes and time series

Movement, demand and risk unfold in time. We model the process, not just snapshots of it.

  • Point processes for arrivals, incidents and events, including self-exciting models where one event raises the rate of the next.
  • State-space models and filters for noisy trajectories, with explicit treatment of missing and irregular observations.
  • Change-point and regime detection for systems whose behaviour shifts.
  • Hazard and survival models for durations: how long until a vehicle moves, a customer leaves, a component fails.

In practice: a session of vehicle pings is a trajectory with gaps. The gap length, the speed on either side of it and the road class at both ends carry information about whether the vehicle stopped or the sensor did.

Graph theory and network methods

Many hard problems are linkage problems: which observations belong together, how flow moves across a network, where the communities are.

  • Association and record linkage under rotating or missing identifiers, framed as assignment problems with calibrated scores.
  • Flows on road, supply and communication networks; origin and destination estimation from partial observation.
  • Community detection and structural analysis of large sparse graphs.
  • Candidate generation under space-time constraints so that linkage scales to millions of nodes.

In practice: at central-London density a single journey end has around 120 plausible continuations. The graph methods decide which to score; the likelihood decides which to keep.

Geospatial statistics

Location data has its own statistics: projection, map-matching, spatial correlation and the geometry of the road network all shape what a model can see.

  • Map-matching of noisy GPS to the road network with a routing engine, and route likelihoods that respect turn restrictions and one-way streets.
  • Travel-time priors by road class and hour, learned from observed traversals, to separate congestion from genuine stops.
  • Spatial point processes and kernel methods for density, hotspots and catchments.
  • Geofencing and dwell analysis with explicit uncertainty on boundaries.

In practice: the ratio of observed to free-flow travel time on a link has a characteristic distribution. Where that ratio sits tells us whether a fragment pair is a continuation or a coincidence.

Information theory

Before building a model we ask what the signal can carry. Entropy and mutual information are the tools for that question.

  • Mutual information between observed features and the target, to rank what is worth collecting.
  • Channel-capacity style arguments for how much a sensor or a panel can tell us at a given sampling rate.
  • Privacy-aware aggregation: how much individual information survives an aggregation step, measured rather than asserted.
  • Minimum description length and related criteria for model selection where likelihoods alone mislead.

In practice: a probe panel that rotates identifiers every few minutes has a measurable information ceiling. Knowing it before procurement saves the client from buying data that cannot answer the question.

Calibration and validation

A model is only useful when its confidence means something. We treat validation as the main event.

  • Forced-choice and ranking evaluation: can the model pick the true answer from a dense field, and by what margin?
  • Reliability diagrams and proper scoring rules so stated probabilities match observed frequencies.
  • Accept-or-abstain thresholds tuned to the client's actual cost of a wrong answer, not to a round number.
  • Planted ground truth and synthetic injection into real data, so performance is measured where it matters, not on easy cases.

In practice: we report true-answer rank distributions and margins, not a single accuracy figure. A model that is right 43 percent of the time on first pick and 80 percent within its top three is a different tool from one that is right 43 percent of the time, full stop.

Contact

Bring us a problem that resists the obvious model.

We take on a small number of engagements at a time. Tell us what you are measuring, what you need to decide, and what the data looks like today. We reply within two working days.

Email
info@advcipher.com

Office
IFZA, Dubai, United Arab Emirates
Working with clients in the UK, Europe and the Gulf.

What to include
The decision you need to support, the data sources you have (or want), timescales, and any constraints on privacy or deployment.

Start a conversation