3 comments

  • abhinavexists 27 minutes ago

    Hey HN! Today we’re releasing interfaze-1-lite, an open-weight model for tasks with a deterministic answer: OCR, speech-to-text with speaker diarization, classification, structured extraction and object detection.

    The problem we had been facing: when the consumer of a model’s output is not a human but another piece of code, having a 95% accurate model was not enough to automate things without knowing which 5% were correct. Chat models cannot do this, hence someone needs to verify the answers anyway.

    Lite is a reasoning core plus a collection of specialist models (document reader, line detector, speech recognizer, segmentation). The core chooses the specialists to be executed and produces the answer in your JSON schema.

    The specialist models give you: bounding boxes, word timestamps and a confidence score per each line and word. The raw outputs are returned next to the answer so that you can filter based on that. E.g., post rows above 0.9 automatically and forward the rest to a human.

    Benchmarks against Gemini-3.7-Flash, Claude-Sonnet-5, GPT-5.4-Mini and Grok-4.3, where Lite does the best: - MMMU-Pro: 73.2% - RefCOCO: 83.8% - Structured output (SOB): 81.5% - olmOCR: 83.8% - VoxPopuli WER: 3.0%

    Runs on a single 80 GB GPU (e.g. H100) without any additional services.

    Blog post: https://interfaze.ai/blog/the-first-open-weight-model-for-de...

    Happy to answer any questions regarding the architecture, benchmarks and limitations.

  • an hour ago
    [deleted]
  • an hour ago
    [deleted]