Skip to content
Made with Jev

Article · by ggml-org

Decision models in llama.cpp

llama.cpp server now answers on /v1/systemone.

The llama.cpp server now supports decision models through a /v1/systemone endpoint. You send a state and typed questions, and the model returns a probability for each option in one forward pass. The API follows the System One format that came with Jev, so an existing client only needs a new base URL. The supported models at launch are Julia-1, Laya, Kev-4B, lev, OpenJev and Clef, with median times from 3 ms to 43 ms for one question on one NVIDIA RTX PRO 6000.

Open the source

More like this