One street photo in, a structured observation record out.
OddEye 0.1 27B is the street-reading model of michiyomi (experimental). It writes what it sees (observations) and what it infers (interpretations) in separate fields, in Japanese. It is compressed to 4-bit, so it runs on one GPU.
- Apache-2.0
- Runs on one RTX 5090
- Released 2026-10-03, experimental
Two eyes: observation and interpretation
In the name, the round parts of the two letters d are two eyes in different colors. Blue stands for observation and amber for interpretation.
What is visible in the photo: objects, text on signs, relations between objects, surface conditions and so on.
What it infers from the photo. Each one carries a confidence label, in a field separate from the observations.
The record has 13 top-level fields, including objects, text on signs, relations, surfaces, standard road facets (lanes, sidewalks, lighting and so on), interpretations and a scene summary.
Comparison results
Outputs of the original model (BF16), a generic 4-bit version and OddEye were compared on 100 street photos without knowing which model wrote which, using the recommended prompt (v5.2c).
| Measure | Original (BF16) | Generic 4-bit | OddEye |
|---|---|---|---|
| Reference facts captured (0 to 1)↑ higher is better | 0.54 | 0.53 | 0.53 |
| Reference facts contradicted↓ lower is better | 9.3% | 12.4% | 11.3% |
| Overconfident statements per photo↓ | 0.70 | 0.97 | 0.69 |
| Clear errors per photo↓ | 0.35 | 0.43 | 0.26 |
| Useful observations beyond the references, per photo↑ | 1.67 | 1.58 | 1.72 |
How to read it
- Compared with the generic 4-bit version, OddEye made 0.28 fewer overconfident statements per photo [−0.44, −0.12] and 0.17 fewer clear errors per photo [−0.31, −0.03].
- Compared with the original model, there was no significant difference in overconfident statements or clear errors. Contradictions with the reference facts, however, were 2.0 points higher [0.2, 3.6], and 3.4 points higher on the 40 photos never seen during development. The generic 4-bit version shows the same rise, so this mainly comes from compressing to 4 bits.
- By kind of fact, relations between objects are the best of the three (score +0.053 and contradictions −4.8 points against the generic 4-bit version). Text is better than the generic 4-bit version (+0.054) but does not reach the original model. For basic layout such as lane counts and sidewalks, contradictions are 5.4 points higher than the original model.
Weak points
- It writes more generic, low-information interpretations than the original model (+0.27 per photo).
- Basic layout (lane counts, sidewalks) and small text are weaker than in the original model. Check those fields when they matter.
- It transcribed a license-plate number in 1 photo of 140 (so did the original model).
How it was evaluated
- 100 street photos from Japan, 40 of which were never seen while developing the prompt or the model. None share an image sequence with the training photos.
- Each photo has about 15 reference facts written by AI from the photo (not checked by people). A Claude model (Anthropic) graded all outputs side by side without knowing which model wrote which.
- The 60 photos used in development score higher for every model (0.47 to 0.49 on the 40 unseen photos, 0.56 to 0.58 on the 60). Read the numbers as a comparison between models, not as absolute accuracy.
Speed and hardware
| Measure | Original (BF16) | Generic 4-bit | OddEye |
|---|---|---|---|
| Weights | 55.6 GB | 20.6 GB | 20.6 GB |
| Images per hour (RTX PRO 6000, 32 at once) | 1,020 | 1,749 | 1,833 |
| RTX 5090 (32 GB) | Does not fit | Not tested | Median 27.5 s per photo, about 900 images per hour with 8 at once |
- On an RTX PRO 6000 it is about 1.8 times as fast as the original model.
- On an RTX 5090, all 13 fields were present in all 29 photos tried. In the 140-photo evaluation, 98.6% of outputs had all 13 fields.
- About 1 run in 70 is cut off at the 8,192-token output limit.
- Speeds were measured on our own evaluation images. Other images or settings will differ.
How it was built
- The base model is Qwen3.8-27B (Apache-2.0).
- It was trained on 400 answers written by the base model itself (BF16) as targets, while simulating 4-bit error (4-bit-aware self-distillation), then calibrated to michiyomi's prompt.
- No outputs of external AI services were used for training.
- The training photos are street photos from Mapillary (not included in the model repository).
How to run it
Start it with vLLM and send photos with the bundled example_request.py (Python standard library only). Use the bundled prompt prompt_v5.2c.md. The model was specialized on it.
Tested with vLLM 0.27 and 0.28 on Blackwell GPUs.
Before you use the output
- It describes only what is visible in that photo, at the time it was taken, not the current state of the place.
- Widths and lane counts are estimates from one image, not measurements. Not visible is a real answer. Do not fill in what the record does not say.
- Redact license plates and personal names before publishing outputs.
- Do not use it as the only basis for safety or accessibility decisions.
Version and license
- Version
- 0.1 (experimental), released for the Final Stage of the Tokyo Governor's Cup 2026. Later versions will be separate repositories, and the full version, OddEye 1.0, will be built on the next Qwen generation.
- License
- Apache-2.0 (the same as the base model). No conditions on the outputs.
- Size
- 20.6 GB in NVFP4 (4-bit floating point).