NYC Mobility Operations

From 3.7 million trips
to one operating decision

An explainable review system for taxi operations. Every alert ends with evidence, an investigation prompt, and a metric to test.

Decision-support prototype January 2026 No production impact claimed
01 · Before and after

From data burden to operating decision

Ready
Before
0 trip records
Raw rows
Manual search
No priority
FDE build
Scope Clean Rank Prove
After
0 ranked prompts
Evidence
Priority
Human decision
94.8%fewer alerts
436 / 436evidence-linked
8 / 8tests passed
Nextfleet impact trial
02 · Signal to action

Three operating decisions

Select a case to follow the evidence chain.

Where the signal appeared

Early-January alert geography

Top alert zones plus the Times Square trip-volume deviation.

176 alerts across 45 zones
NYC taxi zone Scenario focus

Hover or focus a highlighted zone.

Geometry: official NYC TLC taxi zones, February 2026 snapshot.

Select a highlighted zone Alert evidence appears here.
Loading official TLC zones
Observed signal
Trip-volume concentration 176 alerts · 00:00–05:00
Market response
Demand moves outward
Supply stays inside plan
Mismatch cars in the wrong zones
Human decision Review operating context before action
Impact to test
Moneymissed trips
Timeempty travel
Riskpickup friction

Trip-volume deviations concentrated in several zones. An operator would inspect source records and local context before deciding whether any staging test is justified.

Verified signal Hypothesis to validate Impact requires fleet trial

Measured and unmeasured impact

Proved now

System impact

  • 94.8%fewer alerts than the first configuration
  • 436 / 436alerts linked to evidence receipts
  • 8 / 8deterministic tests passed
Requires fleet trial

Business impact

  • Tripscompleted per driver-hour
  • Milesdriven without a passenger
  • Timepassenger pickup and airport waiting
03 · Use and test

Move from map signal to proof

Use the report
  1. 1
    Select an alert patternConcentration, airport, or duration shift
  2. 2
    Inspect highlighted zonesFind the place and hour
  3. 3
    Review the evidenceSignal before explanation
  4. 4
    Choose whether to investigateHuman decision required
Test the impact
  1. A
    Choose test zonesFive alert-informed zones
  2. B
    Choose control zonesFive comparable zones
  3. C
    Change one decisionStaging, dispatch, or ETA
  4. D
    Measure the differenceMoney, time, risk, service
Trips per driver-hour Revenue per online hour Empty miles Pickup delay
The “so what?”

It turns anomaly detection into an operating test.

Signal → evidence → human decision → controlled action → money and service metric.

Human review dashboard

436 prompts.
One review queue.

A bounded January 2026 snapshot of the final anomaly queue. Filter first. Investigate second. Do not infer cause from a deviation alone.

Static snapshotOfficial TLC sourceInvestigation only
Review prompts436final queue
High priority343requires first review
Evidence coverage436 / 436linked receipts
Data span31 daysJanuary 2026
Queue explorer

Find a reviewable signal

Loading verified queue snapshot.

Prompts by detector

Highest-volume review dates

PriorityDate and hourZoneDetectorObserved vs baselineReview prompt

Investigation prompt only. This alert does not establish cause, fraud, fault, or a prescribed action.

Interview-ready case brief

Problem.
Process.
Payoff.

A direct explanation of what the system changed, what it proved, and what still needs a fleet trial.

Official TLC data Deterministic methods Business impact not yet claimed
P1
Problem

Data existed. Operating priority did not.

January taxi records described completed trips. They did not tell a fleet operator which zone-hour changes deserved review first.

3,724,889trip records
8,423first-pass alerts
Unrankedoperating queue
The bottleneck was decision overload, not missing data.
P2
Process

Turn raw activity into review-ready evidence.

  1. 1
    Quality gateType records, quarantine rejects, document null policy
  2. 2
    Comparable baselineCompare each zone by weekday and hour
  3. 3
    Explainable detectionVolume, fare, duration, fare-per-mile, payment mix
  4. 4
    Evidence receiptLink every prompt to the supporting Gold-table query
  5. 5
    Human boundaryPrompt investigation. Do not prescribe action or claim cause
The FDE work connected data quality, model logic, operator workflow, and evidence.
P3
Payoff

Smaller queue. Stronger proof. Review ready.

94.8%fewer alerts than the first configuration
436 / 436final prompts linked to evidence
8 / 8deterministic tests passed
ProvedReview system reliability
Not yet provedFleet revenue or service improvement
The current payoff is a smaller, auditable decision queue. Operational ROI needs a controlled pilot.
Real-world deployment path

An internal shift-planning product for a fleet operator.

Proposed pilot · not deployed
Economic buyer Director of Fleet Operations

Owns driver productivity, service levels, and the pilot budget.

Daily users Dispatch and operations managers

Review ranked zone-hour prompts before each operating shift.

Delivery team FDE + data engineer + ops analyst

Connect fleet data, tune thresholds, train users, and measure outcomes.

01Ingest

Combine completed trips with live fleet supply, demand, ETA, and weather data.

02Prioritize

Rank zone-hour changes and attach an evidence receipt to each prompt.

03Approve

An operator accepts, rejects, or annotates a staging or dispatch test.

04Measure

Compare test and control zones, then retain only changes linked to better results.

Decision changed Where and when to stage available drivers
30-day pilot measures Trips per driver-hour · revenue per online hour · empty miles · pickup delay
Ship / stop rule Expand only after test zones outperform matched controls without harming service or safety
With more time

Move from decision support to measured operating impact.

Next approval gate
01Interview operators

Validate which alerts change staging, dispatch, or ETA decisions.

02Run a controlled pilot

Compare five alert-informed zones with five similar control zones.

03Log decisions

Connect signal, approved action, operator rationale, and observed result.

04Measure ROI

Track trips per driver-hour, revenue per online hour, empty miles, and pickup delay.

Build order Operator workflow first → measurement second → narrative or LLM layer last
90-second answer

What did you build, and so what?

“I turned 3.7 million NYC taxi records into 436 evidence-linked investigation prompts. I built the quality layer, robust zone-hour baselines, five explainable detectors, evidence receipts, and a human review boundary. The system reduced alert overload by 94.8% and passed eight deterministic tests. The proposed buyer is a director of fleet operations, with dispatch managers reviewing prompts before shifts. I have not claimed fleet ROI. The next step is a 30-day controlled pilot measuring trips per driver-hour, revenue per online hour, empty miles, and pickup delay.”