Article · by ggml-org
Decision models in llama.cpp
llama.cpp server now answers on /v1/systemone.
The llama.cpp server now supports decision models through a /v1/systemone endpoint. You send a state and typed questions, and the model returns a probability for each option in one forward pass. The API follows the System One format that came with Jev, so an existing client only needs a new base URL. The supported models at launch are Julia-1, Laya, Kev-4B, lev, OpenJev and Clef, with median times from 3 ms to 43 ms for one question on one NVIDIA RTX PRO 6000.
Open the source