The hidden bottleneck
Walk-through facial recognition promises faster, hands-free station access. But nominal scan speed tells only half the story. A failed authentication takes longer than a successful passage, may trigger several immediate retries, and can end with a passenger being sent to staff. At high demand, those failures—not the routine successes—can generate most of the biometric workload.
The speed of a successful scan is only part of the capacity story. Failures determine how much recovery capacity the station needs.
One policy, three coupled queues
The paper models biometric gates, conventional gates, and an exception desk as one connected service system. Passenger groups may differ in biometric adoption, image-acquisition quality, and sensitivity to delay. The recognition threshold and retry limit therefore do more than change accuracy: they determine gate occupation, exception arrivals, waiting, and whether each part of the system remains stable.
This makes exception demand endogenous. Tightening the recognition threshold can reduce false accepts, but it can also create more false non-matches, retries, and staff interventions. Allowing more retries can keep passengers away from the exception desk, yet block biometric gates for longer.
Four decisions that must move together
- Recognition threshold: balance security risk against false non-matches, delay, and recovery workload.
- Retry rule: use a finite limit that reflects actual field reliability; retries move congestion between the gates and the exception desk.
- Capacity: add gates when ordinary throughput is the bottleneck, but strengthen exception service when authentication failures dominate.
- Adoption: stage biometric use with the infrastructure mix. Moving everyone to the biometric channel can increase delay if gate and recovery capacity do not grow with it.
No universal retry count
The numerical results make the need for calibration concrete. In the stylized peak-period experiment, an intermediate threshold and a three-attempt limit minimized the modeled cost. In the public-data-anchored counterfactual, low benchmark error rates made additional retries comparatively inexpensive and moved the preferred retry limit to the top of the tested range. The contrast is the practical result: retry policy should be set from observed acquisition quality and passenger experience, not copied as a universal number.
Fairness beyond average waiting time
Average delay can hide an important disparity. In the paper's stress test, passengers with more difficult image acquisition were routed to exception handling more than five times as often as the regular class at the most severe tested setting, even though the difference in average delay looked comparatively modest. Monitoring who is repeatedly stopped or redirected is therefore as important as monitoring the mean queue.
From a model to station-specific decisions
The study anchors its counterfactuals in 44,168 station-hour observations from New York's subway system and biometric operating points from NIST benchmarks. Under the stated baseline assumptions, extra biometric-gate capacity was valuable at the busiest hubs, including Grand Central–42 St and Times Square, but not at the lower-demand stations. The analytical approximation was also checked against discrete-event simulation: it reproduced the main utilization, delay, and stability patterns, while supporting a conservative safety margin near saturation.
Operational takeaway
Evidence boundary
This page independently summarizes the authors' COMOSA 2026 conference paper. Its stylized experiments and public-data-anchored counterfactual are not an empirical evaluation of a live biometric deployment. MTA ridership data are not biometric-gate logs, OMNY share is only a digital-payment adoption proxy, and the biometric operating points come from NIST benchmarks rather than measurements at the modeled stations.
The numerical findings are scenario-dependent, not universal deployment prescriptions. The study analyzes queueing and operational consequences; privacy, consent, security, accessibility, and governance require separate evaluation.