Whitepaper

michiyomi Design Principles

White paper · 2026-08-28 edition. For usage, see the API wiki.

The 2026-08-28 edition text, word for word. Links now point to the matching places in the new docs, and numbers that changed since carry a dated note (Note (2026-10-03)).

1. Abstract

michiyomi is an open-data platform that converts citizen-published Mapillary street imagery into structured text with vision-language models (VLMs), making physical street conditions searchable by coordinate. It has processed 1,914,490 scenes across all 23 of Tokyo’s wards (1,914,451 searchable); its 2,522 longitudinal-change records include only claims that survived falsification-oriented adjudication. REST and MCP provide unauthenticated, read-only access; observation data is under CC BY 4.0. Note (2026-10-03): release 2026-09-13-r1 covers all of Tokyo: the 23 wards, the Tama area and the islands.

One policy runs through the design: beyond maximizing generation accuracy, build provenance, verification, and limitations into the system itself. We believe that is essential if generative-AI output is to become citable public data.

2. Background and problem

Maps know the shape of a road, but rarely its visible condition: worn lane markings, sidewalk width, tactile paving, blocked sightlines, and how these change over time. For local roads—which make up most of the network—understanding those conditions still often requires a site visit.

The raw material already exists. Mapillary contains vast quantities of street imagery published by citizens and organizations, and VLMs can now read it. What is missing is a design that makes the results trustworthy enough to cite. A fluent AI description does not, by itself, make clear where fact ends and interpretation begins, where the data came from, or how to represent something outside the frame. Administrative triage and research require those ambiguities to be resolved structurally.

People are not the only clients. If geospatial AI agents can ask “what did this place look like?”, verbalized street data can support dialog, navigation, and analysis. The API was therefore designed from the outset with AI agents as first-class users.

3. Design principles

The implementation follows five principles.

3.1Separate observation, computation, and interpretation

Each scene is delivered as three distinct layers: observed facts (capture year, position, and bearing metadata); deterministic features (travel bearing and image-color analysis, reproducible from the same input); and model interpretation (the VLM verbalization). The layers live in separate namespaces and never overwrite one another.

The purpose is to make the boundary between fact and interpretation visible in the data structure, rather than relying on the reader’s vigilance. This enables auditing and internal cross-checks. For example, a VLM’s Japanese description of continuous roadside trees can be compared with the machine layer’s green-view ratio. Every description is traceable through its image ID to the source photograph.

3.2Implement honesty as a specification

The most dangerous mistake is treating a location with no data as a location with nothing there. To prevent it, the coverage endpoint reports the available evidence before content is requested. The MCP tool definitions encode the workflow itself: check coverage first, then retrieve descriptions.

Examples of specified honesty
  • The generation prompt must distinguish absent, unknown, and outside the frame; it must not infer what is not visible.
  • Capture timestamps are not accepted uncritically. After encountering a daylight image timestamped near 23:00, the visual estimate became authoritative for time-of-day description.
  • Thirty-nine anomalous-timestamp scenes are quarantined from search, and both the quarantine and its count are disclosed through the API.

The user-facing rules are consolidated in Important limitations.

When honesty is part of the specification, automated tests and consistency checks can continuously verify whether the implementation still honors it.

3.3Record generation provenance for every scene

A verbalization is generated output, so a model change can change the result. Every one of the 1,914,490 scenes therefore records which model family produced it, under which output contract, and when. This provenance bundle is called a generation.

Generation is an origin label, not a quality ranking. The cloud-VLM gen1 (77,676 scenes) and home-GPU local-VLM gen2 (1,836,814 scenes, thirty-nine quarantined from search) are both adopted for serving. The local generation contains multiple models and inference settings; equivalent judgment criteria across all of those configurations have not been established.

Model migration becomes an explicit governance event. A new family may run alongside existing output, pass the gate, and then join the released set. Earlier output is retained rather than discarded or overwritten. Local-VLM counterparts also exist for gen1-selected scenes, enabling generation-to-generation comparison.

3.4Require falsification-oriented adjudication for change claims

Longitudinal change is among the most error-prone forms of verbalization. A viewpoint shift of a few meters or a seasonal difference can make an unchanged object appear to have changed. We therefore insert adjudication between detection and publication.

First, the system selects only image pairs from the same road group with compatible viewing directions. A VLM describes the apparent difference. Then a separate, context-isolated session is asked to falsify that claim. Only claims that survive as supported are published: currently 3,618. In October 2026 the pairs were re-selected so that photos match by position and heading and re-compared with GPT-6.1 Sol; 2,963 of the 3,025 claims put to the checker survived. High-confidence claims from the August 2026 set (gpt-5.6-sol, first pass) were re-judged blind by the newer model on the same photos, and only the 655 that survived both are kept. About 20% of the 2,522 August claims previously published turned out not to match the photos in this review and were removed. Rejected findings are not included.

Limitations ship with the data. The adjudication is a falsification-oriented re-evaluation by a context-separated session from the same model family, not independent third-party verification. Shared systematic error may remain. The supported rate is the fraction that survived attempted falsification, not an accuracy score or lower-bound guarantee.

3.5Compute geometry deterministically; do not ask the VLM to guess

We want data capable of saying “a fire hydrant is 30 m northeast.” But north is not visible in an image, so the VLM should not infer it.

The VLM reports only image-relative position, such as “a fire hydrant in the right foreground.” Travel bearing is computed geometrically from neighboring frames in the capture sequence and combined with the corrected camera bearing to derive absolute eight-way directions. The computation is deterministic and available for 99.7% of scenes. Note (2026-10-03): 99.6% in release 2026-09-13-r1 (travel_bearing in /v1/meta).

previouscurrentnext travel_bearing = bearing(previous→next) corrected camera bearing offset classifies forward / side / rear views travel bearing ±90° maps image “left” and “right” to eight compass directions
Travel bearing is calculated from preceding and following frame coordinates; nearby stationary frames are excluded.

The same reasoning excludes momentary counts such as “three pedestrians” at prompt time. Descriptions focus on equipment, structures, surfaces, and other features that are more stable over time. Choosing what not to generate extends the useful life of the data.

4. Architecture

4.1Four verbalization layers

A single frame cannot resolve every question. The pipeline therefore organizes verbalization into four levels.

LevelProblem addressed
L1 point — 1,914,490 scenesPhysical state in one image. Input is the image alone: no place name, map, earlier output, or other external knowledge.
L2 segment — about 100 m of sequential framesResolve occlusion, distinguish moving from stationary objects, and understand the extent of damage.
L3 time — same place × different yearDescribe differences between two views, then send claims to the adjudication in principle 3.4.
L4 context — public-data joinsJoin schools, zoning, and other official data in the database without showing them to the VLM, preserving accountability between “what AI saw” and mapped facts.

Keeping external knowledge out of L1 is important: an image-only description remains an independent source when later compared with official data.

4.2Releases and delivery

Data is not published immediately after generation. It is frozen into a versioned release and promoted through raw SQLite, a production-equivalent preview, and production. The same 96 consistency checks—counts, orphan records, coordinate ranges, duplicates, and statistic agreement—run in all three environments. The API is covered by automated regression tests. Every response carries the release ID as snapshot, so a query can later be reproduced against the same version.

Delivery runs on Cloudflare Workers and two shards of the SQLite-based distributed D1 database. Note (2026-10-03): four shards since release 2026-09-13-r1. REST and MCP share the same database, quality contract, read-only posture, and rate limits. Operational lessons are codified too: after a missing ANALYZE step degraded geospatial queries to tens of seconds, mandatory post-import analysis became part of the runbook.

5. Quality assurance and known limitations

Quality assurance rests on four reproducible controls: (1) regression-comparison and calibration gates for generations, (2) 96 consistency checks across three environments, (3) automated API regression tests, and (4) falsification-oriented adjudication for longitudinal changes.

The following limitations are permanent parts of API responses, tool definitions, and public documentation—not footnotes.

LimitationImplication
Descriptions are estimates from a captured imageThey do not establish current conditions or safety. Widths are monocular estimates, not survey measurements.
Adjudication is not independent verificationIt uses context-separated sessions from the same model family; family-wide errors can remain.
Coverage depends on available imageryAn unphotographed road is absent from the dataset, so coverage should be disclosed before interpretation.
Judgments differ across models and inference settingsQwen3.6 and 3.8 show large differences in absent versus out-of-frame sidewalk labels. Their input populations also differ, so the cause is unconfirmed. Matched-image evaluation is needed before comparing regions or years.
Color features depend on capture conditionsGreen-view and related metrics include camera and lighting effects; aggregate analysis should stratify by capture conditions.

6. Operating model

michiyomi is a personal project designed to remain operable after the event.

  • Delivery: Edge and serverless infrastructure avoids ongoing server administration.
  • Generation: The main L1 pipeline uses local VLMs on home GPUs. Initial L1 records, L2, L3 and adjudication have used cloud models. Local L2 processing is under evaluation.
  • Updates: New data enters through the existing release and generation mechanisms, making incremental updates routine.

Unauthenticated access is deliberate. People and AI agents should be able to submit coordinates and retrieve street evidence without registration. Rate limiting and read-only design provide the guardrails.

7. Roadmap

At this edition, the platform has completed the target scene set across all of Tokyo: 1,914,490 processed scenes and 3,618 adjudicated changes. Coverage density still reflects the uneven distribution of source Mapillary imagery.

Short term
through Aug 2026
The 23-ward target set was completed in Aug 2026; release 2026-09-13-r1 added the Tama area, the islands (latest year) and the 23-ward past years, reaching 1,914,490 scenes across all of Tokyo. Verbalization, machine-layer processing, and release preparation reached 1,009,250 scenes. Note (2026-10-03): 1,009,250 is the 23-ward count of release 2026-08-28-r1. Subsequent imagery increments will be published as ordinary releases; nationwide expansion and later model generations will use geographic shards and generation archives.
Medium term
within 2026
1. Unify the served generation. Local-VLM verbalization is already complete for all scenes; after adoption gates, the served set will converge on one generation family while retaining older provenance.
2. Bulk data distribution. A public Parquet snapshot is available on Hugging Face for research and analysis (observation data under CC BY 4.0; see License). Its update time may differ from the live API.
3. Extend the time axis. Expand road-level change visualization and longitudinal coverage using multi-year imagery.
Long term
1. Connect to public-sector workflows. Support inspection prioritization and pre-visit triage—“read before going”—through a low-integration-cost, coordinate-addressable read-only API.
2. Expand geographically. The pipeline has no Tokyo-specific architectural requirement and can be applied wherever open street imagery exists.
3. Establish an update cycle. Reduce the time from citizen capture to searchable description, and test wearable uses such as real-time pedestrian audio guidance.

8. Conclusion

michiyomi is intended to demonstrate not only individual verbalization accuracy, but a way to circulate generative-AI output as public data: three-layer separation, specified honesty, generations as provenance, falsification-oriented adjudication, and deterministic division of labor. These are data-platform disciplines independent of any particular model.

Models will continue to change. The center of this project is building the data platform so that those changes do not break its auditability.

michiyomi Design Principles white paper (2026-08-28 edition) · API wiki · Development history · Home

Photos: © Mapillary contributors (CC BY-SA 4.0) / observation data: michiyomi, CC BY 4.0 / school locations based on MLIT National Land Numerical Information (P29). See License.

Location descriptions are VLM estimates from imagery captured in the stated year; they do not determine current conditions or safety.