Model
OddEye

One street photo in, a structured observation record out.

OddEye 0.1 27B is the street-reading model of michiyomi (experimental). It writes what it sees (observations) and what it infers (interpretations) in separate fields, in Japanese. It is compressed to 4-bit, so it runs on one GPU.

  • Apache-2.0
  • Runs on one RTX 5090
  • Released 2026-10-03, experimental

Model: huggingface.co/finalvent/OddEye-0.1-27B-NVFP4. The numbers on this page are the evaluation results published on the model card.

The public data on this site was not made with OddEye. The verbalizations served by the API, MCP and the dataset come from the production pipeline (its current prompt is the v3.5a family). OddEye is a separate model for a next step that writes richer observations (v5.2, 13 fields). How the public data is made has not changed yet.

Two eyes: observation and interpretation

In the name, the round parts of the two letters d are two eyes in different colors. Blue stands for observation and amber for interpretation.

Observation

What is visible in the photo: objects, text on signs, relations between objects, surface conditions and so on.

Interpretation

What it infers from the photo. Each one carries a confidence label, in a field separate from the observations.

The record has 13 top-level fields, including objects, text on signs, relations, surfaces, standard road facets (lanes, sidewalks, lighting and so on), interpretations and a scene summary.

Comparison results

Outputs of the original model (BF16), a generic 4-bit version and OddEye were compared on 100 street photos without knowing which model wrote which, using the recommended prompt (v5.2c).

MeasureOriginal (BF16)Generic 4-bitOddEye
Reference facts captured (0 to 1)↑ higher is better0.540.530.53
Reference facts contradicted↓ lower is better9.3%12.4%11.3%
Overconfident statements per photo↓0.700.970.69
Clear errors per photo↓0.350.430.26
Useful observations beyond the references, per photo↑1.671.581.72

How to read it

  • Compared with the generic 4-bit version, OddEye made 0.28 fewer overconfident statements per photo [−0.44, −0.12] and 0.17 fewer clear errors per photo [−0.31, −0.03].
  • Compared with the original model, there was no significant difference in overconfident statements or clear errors. Contradictions with the reference facts, however, were 2.0 points higher [0.2, 3.6], and 3.4 points higher on the 40 photos never seen during development. The generic 4-bit version shows the same rise, so this mainly comes from compressing to 4 bits.
  • By kind of fact, relations between objects are the best of the three (score +0.053 and contradictions −4.8 points against the generic 4-bit version). Text is better than the generic 4-bit version (+0.054) but does not reach the original model. For basic layout such as lane counts and sidewalks, contradictions are 5.4 points higher than the original model.

Weak points

  • It writes more generic, low-information interpretations than the original model (+0.27 per photo).
  • Basic layout (lane counts, sidewalks) and small text are weaker than in the original model. Check those fields when they matter.
  • It transcribed a license-plate number in 1 photo of 140 (so did the original model).

How it was evaluated

  • 100 street photos from Japan, 40 of which were never seen while developing the prompt or the model. None share an image sequence with the training photos.
  • Each photo has about 15 reference facts written by AI from the photo (not checked by people). A Claude model (Anthropic) graded all outputs side by side without knowing which model wrote which.
  • The 60 photos used in development score higher for every model (0.47 to 0.49 on the 40 unseen photos, 0.56 to 0.58 on the 60). Read the numbers as a comparison between models, not as absolute accuracy.

Speed and hardware

MeasureOriginal (BF16)Generic 4-bitOddEye
Weights55.6 GB20.6 GB20.6 GB
Images per hour (RTX PRO 6000, 32 at once)1,0201,7491,833
RTX 5090 (32 GB)Does not fitNot testedMedian 27.5 s per photo, about 900 images per hour with 8 at once
  • On an RTX PRO 6000 it is about 1.8 times as fast as the original model.
  • On an RTX 5090, all 13 fields were present in all 29 photos tried. In the 140-photo evaluation, 98.6% of outputs had all 13 fields.
  • About 1 run in 70 is cut off at the 8,192-token output limit.
  • Speeds were measured on our own evaluation images. Other images or settings will differ.

How it was built

  • The base model is Qwen3.8-27B (Apache-2.0).
  • It was trained on 400 answers written by the base model itself (BF16) as targets, while simulating 4-bit error (4-bit-aware self-distillation), then calibrated to michiyomi's prompt.
  • No outputs of external AI services were used for training.
  • The training photos are street photos from Mapillary (not included in the model repository).

How to run it

Start it with vLLM and send photos with the bundled example_request.py (Python standard library only). Use the bundled prompt prompt_v5.2c.md. The model was specialized on it.

vllm serve finalvent/OddEye-0.1-27B-NVFP4 --served-model-name oddeye \ --max-model-len 16384 --kv-cache-dtype fp8 --mamba-ssm-cache-dtype bfloat16 \ --limit-mm-per-prompt '{"image":1}' \ --speculative-config '{"method":"mtp","num_speculative_tokens":2}' --enable-prefix-caching # RTX 5090 (32 GB): add --max-num-seqs 8 --gpu-memory-utilization 0.92 python3 example_request.py photo.jpg > observation.json

Tested with vLLM 0.27 and 0.28 on Blackwell GPUs.

Before you use the output

  • It describes only what is visible in that photo, at the time it was taken, not the current state of the place.
  • Widths and lane counts are estimates from one image, not measurements. Not visible is a real answer. Do not fill in what the record does not say.
  • Redact license plates and personal names before publishing outputs.
  • Do not use it as the only basis for safety or accessibility decisions.

Version and license

Version
0.1 (experimental), released for the Final Stage of the Tokyo Governor's Cup 2026. Later versions will be separate repositories, and the full version, OddEye 1.0, will be built on the next Qwen generation.
License
Apache-2.0 (the same as the base model). No conditions on the outputs.
Size
20.6 GB in NVFP4 (4-bit floating point).

Read the model card on Hugging Face ↗