Live site · by r33drichards
Laya Vision
Typed decisions about an image plus text, in one forward pass.

An independent fork of Laya that answers choice, yes/no and score questions about a picture and optional text. The checkpoint is SmolVLM-256M-Instruct cut to 20 of its 30 language layers, trained for 2 hours on question-answering data and game frames. An autoresearch loop found the configuration.
Open the source