Mining open datasets for transparency in taxi transport in metropolitan environments

Noulas, Anastasios and Salnikov, Vsevolod and Lambiotte, Renaud and Mascolo, Cecilia (2015) Mining open datasets for transparency in taxi transport in metropolitan environments. EPJ Data Science, 4.

[thumbnail of s13688-015-0060-2]
PDF (s13688-015-0060-2)
s13688_015_0060_2.pdf - Published Version
Available under License Creative Commons Attribution.

Download (3MB)


Uber has recently been introducing novel practices in urban taxi transport. Journey prices can change dynamically in almost real time and also vary geographically from one area to another in a city, a strategy known as surge pricing. In this paper, we explore the power of the new generation of open datasets towards understanding the impact of the new disruption technologies that emerge in the area of public transport. With our primary goal being a more transparent economic landscape for urban commuters, we provide a direct price comparison between Uber and the Yellow Cab company in New York. We discover that Uber, despite its lower standard pricing rates, effectively charges higher fares on average, especially during short in length, but frequent in occurrence, taxi journeys. Building on this insight, we develop a smartphone application, OpenStreetCab, that offers a personalized consultation to mobile users on which taxi provider is cheaper for their journey. Almost five months after its launch, the app has attracted more than three thousand users in a single city. Their journey queries have provided additional insights on the potential savings similar technologies can have for urban commuters, with a highlight being that on average, a user in New York saves 6 U.S. Dollars per taxi journey if they pick the cheapest taxi provider. We run extensive experiments to show how Uber’s surge pricing is the driving factor of higher journey prices and therefore higher potential savings for our application’s users. Finally, motivated by the observation that Uber’s surge pricing is occurring more frequently that intuitively expected, we formulate a prediction task where the aim becomes to predict a geographic area’s tendency to surge. Using exogenous to Uber data, in particular Yellow Cab and Foursquare data, we show how it is possible to estimate customer demand within an area, and by extension surge pricing, with high accuracy.

Item Type:
Journal Article
Journal or Publication Title:
EPJ Data Science
Additional Information:
© 2015 Noulas et al. Open Access This article is distributed under the terms of the Creative Commons Attribution 4.0 International License (, which permits unrestricted use, distribution, and reproduction in any medium, provided you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made.
Uncontrolled Keywords:
ID Code:
Deposited By:
Deposited On:
28 Jan 2016 14:06
Last Modified:
18 Sep 2023 00:58