Where the photo record ends and the AI reading begins.
This page tells you which values in a response you can lean on, and how far. michiyomi returns the capture record, deterministic calculations and the AI reading without mixing them, and what could not be determined stays undetermined. The observation text itself is Japanese.
One record is one photo (a scene)
A record, or scene, is one photo. For every 20 m grid square, capture direction (in 45° steps) and capture year, one representative photo is chosen. The same place in a different year or direction is a different scene, and that is what changes over time are compared from.
idis the Mapillary image ID. Even though it looks like a number, keep it as a string. Converting it can lose digits.- Coordinates are where the photo was taken, not where an object stands or where a building entrance is. Mapillary's position corrected from the photo sequence is used by default (1,899,940 of the 1,914,451 published scenes), with raw GPS only when it is missing.
position_sourcesays which. - Of the 1,914,490 processed scenes, 39 with clearly broken capture times are left out of nearby search and counts. They can still be read by ID.
Three layers: record, calculation, AI reading
For each photo, three different kinds of information are kept apart. GET /v1/scenes/{id}?include=metadata,machine,analysis returns all three (the default is analysis only).
metadata · the capture record- Capture date, time and year
- Position (corrected and raw) and camera heading
- Capture sequence and image quality
machine · deterministic- Direction of travel (from neighboring frames)
- Compass direction of objects (north, southeast…)
- Color and green share, season and time-of-day bucket
analysis · the AI reading- Sidewalks, widths, paving, faded markings
- Sign text, buildings, planting
- Risk cues, and what could not be seen
Keeping the layers apart lets you notice disagreements. Photo 964535250779795 (2019, Chuo City) has a capture time of 22:54 in its record, so the calculation derived from it says night. The AI reading says daytime (time_of_day: 日中) and notes backlight washing out the buildings. Capture times depend on the camera clock, so the AI's visual estimate is treated as authoritative for time of day (Whitepaper 3.2).
summary in search results is written for people. For programs, use the structured analysis.Generations: which AI read it
Every description carries a generation (generation): the model family and output contract that wrote it. A generation is a provenance label, not a quality ranking. Each record also carries the model name (model), the prompt version (prompt_version) and when the text was written (generated_at).
| Generation | Which AI read it | Published scenes |
|---|---|---|
gen2-qwen-local (default) | Local VLM (Qwen family, on a home GPU) | 1,836,775 |
gen1-codex | Cloud VLM (GPT family) | 77,676 |
By default, one generation's description is returned per photo. To see one generation only, pass generation=gen1-codex and so on. The catalog is at GET /v1/generations.
Confidence and estimates
Widths, lane counts, distances and similar numbers are estimates from one photo, not measurements. Many fields carry the evidence (evidence) and a confidence (confidence: 高, 中, 低 for high, medium, low). The summary writes [high] [mid] [low]. There is also an overall confidence_overall (0 to 1).
// Roadway width of photo 1499911054492109 (2025, Shinjuku City), from analysis.geometry "roadway_width_m": { "value": 4, "evidence": "駐輪している原付バイク(幅約0.7m)をスケール基準に…", "confidence": "中" }
The evidence says the width was estimated using a parked moped (about 0.7 m wide) as the scale reference (excerpt, translated).
Confidence is the model's own judgement, not a separately measured accuracy. Risk cues (risk_cues) are cues read from the photo, not accident statistics or an official safety assessment.
None, unknown and outside the frame
So that nothing invisible gets filled in by guesswork, the AI is required to keep three answers apart.
In the same Shinjuku photo, the right-hand sidewalk is outside the frame, and road cracks and tactile paving are unknown. What the camera could not see is listed in off_frame and what could not be judged in not_assessable, each with a reason.
"sidewalk": { "left": { "presence": "あり", "type": "路側帯", "width_m": 1.5, "confidence": "中" }, "right": { "presence": "画角外不明", "type": "該当なし", "width_m": null, "confidence": "低" } }, "off_frame": ["道路の反対側(対向側の歩道・沿道)は画角外のため評価不可", …]
The left sidewalk is a roadside strip (路側帯) about 1.5 m wide. The off_frame entry says the opposite side of the road is outside the frame and cannot be assessed.
Do not read unknown as none. When you aggregate, do not count unknown or outside the frame as none. The rates in street and town-block kartes use only the cases that could be judged as the denominator.
When it describes
- A description is what the place looked like in the capture year. Every result has a capture year (
capture_year). Do not present a description of a 2019 photo as the current state. generated_atis when the text was written, not when the photo was taken.data_as_ofis when the release was fixed.- Capture years run from 1986 to 2026, but 99.8% of scenes are from 2014 onward (3,660 scenes predate 2014). Older records, especially those before 2010, are more likely to carry camera-clock errors.
- The 23 wards include photos from several years. The Tama area and the islands are mostly the latest year.
Coverage: only where photos exist
michiyomi is built from photos that people took and published on Mapillary. Where photos exist, and how many, depends on where those people walked and drove, so it varies a lot from place to place.
Check what was recorded with GET /v1/coverage before answering. Within 300 m of the Ginza 4-chome crossing, for example, there are 2,312 photos taken between 2010 and 2026. Zero means nobody photographed there. It does not mean nothing is there, or that nothing changed. When candidate_truncated is true in a nearby search, the candidates were cut off partway (always on the side farther from the center).
Street and town-block kartes
A karte gathers the readings of individual photos into a street (a line) or a town block (an area) and writes them up as text, in Japanese.
- Streets
- OpenStreetMap roads cut at intersections into segments (15 or more photos each), grouped by name within one municipality. Look them up by
street_id. Unnamed roads only have numbers per segment (edge_id). - Town blocks
- 2020 census small areas. Look them up by
town_key. - How it is written
- The text is written only from an aggregated profile (
profile), and a machine check confirms that every number, year and proper noun in it exists in that profile. Nothing outside the profile is written. - Rates
- Shares such as how often there is a sidewalk use only the cases that could be judged as the denominator (
judged, with the judged share injudged_rate). They pool several years and are not the current state. Quoteyear_spanwith them. - Sides
- The two sides of a street are named by compass direction, such as the northeast side. They are built only from forward- and rear-facing photos, so counts can be small.
- Versions
- The model and prompt version behind the text are in
karte_provenance. Versions written by other models are listed inkarte_versions, andinclude=versionsreturns their text too. Every version comes from the same profile and passes the same machine check.
Coverage grades
A rough measure of how many photos there are and how widely they spread. Grade D has no text karte, only numbers via include=profile (the rule is in profile.coverage.rule of the response).
| Grade | Streets (all must hold) | Town blocks (all must hold) | Karte |
|---|---|---|---|
| A | 100+ photos, 5+ capture runs, photos along 60%+ of the length | 300+ photos, 10+ runs, photos along 40%+ of the roadway length | Yes |
| B | 40+ photos, 3+ runs, 40%+ of the length | 100+ photos, 5+ runs, 25%+ of the roadway length | Yes |
| C | 15+ photos, 2+ runs | 30+ photos, 2+ runs | Yes |
| D | Fewer | Fewer | No (numbers only) |
The current version (la_release 2026-10-03-la2) has 6,248 streets (3,563 with a karte), 5,256 town blocks (4,734 with a karte) and 30,362 road segments.
A street where a multi-lane roadway, wider than the ward average, runs northwest, lined on both sides by high-rise office buildings and apartments.Opening of the Harumi-dori (Chuo City) karte, translated from Japanese. Grade A, 1,739 photos from 2013 to 2026. Original: GET /v1/streets/799
Changes over time: only those that held up
Of the claims that something changed, found by comparing photos of the same place from different years, only those that held up against an attempt to refute them are published.
- Pick pairs to comparePhotos from different years along the same road group, facing compatible directions, are paired mechanically.
- Describe the differenceAn AI writes what changed between the two times.
- Try to break itA separate AI session (same model family), with its context cut off, re-examines the claim with the aim of refuting it.
- Publish only what survivedOnly claims that held up (
supported) enter the API.
Breakdown of the 2,660 claims in the published first pass. 94.8% survived.
- Coordinates are the center of the road group, not the position of the changed object.
segment_priority_scoreis a heuristic for inspection priority on a road segment, not the danger of the change itself (it isnullfor many rows).- The IDs of the photos used as evidence (
scene_ids_aandscene_ids_b) are not in the source data, so they arenullin the current release. The full evidence text is atGET /v1/changes/{change_id}.
In 2017 a square red-framed sign with a taxi pictogram, "22-1" and left and right arrows is visible. In 2026 it has become a round taxi-related regulatory sign, with a "22:00-1:00" supplementary sign below it and a supplementary plate with left and right arrows.Taxi-related signs in front of GINZA SIX (equipment, 2017 to 2026), translated from the Japanese evidence. Original: GET /v1/changes/chg_4424424a5b4f
Stability: reading the same photo twice
The same 36,479 photos were read twice by the AI under the same conditions, and these are the shares where both answers matched. They show how steady the readings are, not how correct they are.
- Capture conditions (weather, time of day)95%+
- Road type85–90%
- Sidewalk present or not75%
- Exact width match57%
Width values match exactly in under six of ten repeat reads. Do not decide anything on one photo's number. Read it together with nearby photos and the tendencies in the street or town-block karte.
Amenities (vending machines, toilets, benches)
There is also a lookup for capture points where the AI saw a vending machine, toilet or bench, nearest first (GET /v1/amenities/nearby). Coordinates and distances are where the photo was taken. Where the amenity actually stands, whether it is still there and whether it is usable are not verified. Some may be on private land or meant only for a facility's users.
Five things to do when you use it
- Give the capture year (
capture_year). Do not present it as the current state. - Treat numbers as estimates. Do not state them as fact.
- Zero means not photographed. Check
coveragefirst. - Do not fill in unknowns. Keep none, unknown and outside the frame apart.
- Credit is not required in answers or displays. For redistribution as data, see License and attribution.