The Agent loop is probably what matters (for local LLM)
The commercial ones seemed to want to monopolize the agent loop.
Today the chat completions API is probably a 'defacto' way of talking to the models
https://github.com/ggml-org/llama.cpp/tree/master/tools/server#post-v1completions-openai-compatible-completions-api
https://vercel.com/docs/ai-gateway/sdks-and-apis/openai-chat-completions
btw, credit goes to the origin:
https://developers.openai.com/api/docs/guides/completions
A thing is, more recent efforts seem to be instead offering just an \*agent\* at the API and putting this \*agent\* layer between you and the model.
local LLM will remain \*very\* important because as is currently, you own the agent loop.
You write that "small little" front / stub that is the agent loop talking to the LLM.
it is day and night difference , practically 2 different universes