We Gathered the Data. Nobody Connected the Pipes.
In 2005, TriMet in Portland and Google agreed on a file format for transit schedules. That is the whole origin story of GTFS. One schema definition, applied consistently, and within a few years hundreds of agencies across the world were publishing to it. What came next is the part nobody planned: trip planners, accessibility tools, transit research, service optimization platforms, an entire ecosystem built by people who never had to phone a single transit agency to ask what a column meant. The standard came first. The applications came because the standard existed and anyone could depend on it.
City infrastructure data never had that moment, and living without it is the reason so much civic open data sits published and unused.
The belief underneath all of this is simple enough that cities already agree with it in public. Municipal data should work the way roads do: maintained at public expense, open to everyone, built to carry load from whoever wants to use it. A road that exists on a map but has inconsistent signage, unpublished maintenance schedules, and no standard for connecting to the road next to it is not useful public infrastructure. Neither is a dataset that exists in a portal but cannot join to anything else without a phone call.
Publishing is the part cities have done. Most major cities now run a portal, and the datasets in them are real. What is missing downstream is governance: consistent schemas so two datasets join without manual cleaning, update cadences that are published and honored, version control that tells a developer whether the file they pulled last month still matches today’s, and documentation written for someone who does not work at the city and cannot find the person who created the dataset two administrations ago. Without that layer, an open data portal is a warehouse where nothing has a label.
The Toronto datasets that SolveTO runs on are detailed in a way that still surprises me. I mapped 418,000 pieces of public infrastructure onto one map before I understood what I was looking at. Catch basin records going back to the 1840s. Sewer network topology. Hydrant maintenance history. Flooding study designations. Traffic signal timing. These exist because city engineers have needed them for internal operations across generations, and publishing them was the right call. What using them as a connected system revealed is that the schemas do not always agree with each other, the update cadences are not always documented, and the datasets most relevant to a single infrastructure problem sometimes use geographic identifiers that will not cleanly join to the identifiers in the dataset sitting right beside them. The data is there. The plumbing between the data is not.
There is no schema standard that travels with a catch basin dataset and specifies its update frequency, its coordinate reference system, its identifier scheme, and its relationship to adjacent datasets. Every city publishing infrastructure data makes independent decisions on all four. Every developer who wants to work across more than one city starts over. That is not a technology gap. GTFS was a text file specification, and it moved an industry.
Some places are closing it. The EU Data Act is phasing in, with core provisions applying from 2025 and the rules on cloud interoperability and portability running through 2027, and it treats data sharing and interoperability as obligations rather than aspirations. Helsinki and Espoo publish open 3D models of their built environments that anyone with a browser uses for permit analysis, climate adaptation planning, and service design. Cambridge’s 2026 open data plan puts publication first and then names governance, standardization, and integration into planning workflows as the conditions that make the data worth anything once it is out. San Francisco has run a 311 dataset going back to 2008 on a daily update. New York maintains comprehensive open infrastructure data on published schedules. What those places share is that somebody’s job is the governance layer. The schema stays consistent because there is a standard someone is accountable to. The cadence gets met because missing it would be a public failure rather than an internal oversight.
Everywhere else, publishing a CSV gets described as transparency. It creates the appearance of openness while preserving the friction, and it is a comfortable thing to report at a council meeting because the work of maintaining it never appears on the agenda. This is the same shape as government’s AI problem being a procurement problem: budgets fund launches and nobody owns the decade after. A dataset with no maintenance line is a pilot that was allowed to keep its URL.
Open data called infrastructure implies something that carries load: dependable, maintained, standardized enough that others build on it without asking permission first. Most cities have published data. Fewer have built that, and the difference is not in the dataset. It is in whether maintenance was treated as a civic obligation or as a launch.
The data layer a city publishes is either an invitation or an illusion. An invitation has clean schemas, stable identifiers, published cadences, version history, and documentation that assumes the reader does not work for the city. An illusion is a file drop with no commitment behind it. Most cities have issued the illusion, a few have built the invitation, and the ones that built it are the ones developers keep coming back to. Transit worked out which of the two it wanted to be in 2005, and everything else in the city is still deciding.
Be the first to comment
Thank you! Your comment will appear shortly, usually within a couple of minutes.
Something went wrong. Try again or contact me directly.