URL: ai.bvoy.dev
Open WebUI is a self-hosted AI chat interface — like ChatGPT, but running on the
homelab. It routes your messages through a mix of cloud models and local models
depending on which you pick.
Login: Sign in with your bvoy.dev account via SSO. No separate account
needed.
Pick a model from the dropdown at the top of the chat window. Each conversation
can use a different model.
| Model | Speed | Best for |
|---|---|---|
| Gemini 3 Flash | Fast | Default — good for most things |
| DeepSeek V4 Flash | Fast | Code and technical questions |
| Claude 3 Haiku | Fast | Writing and summarization |
| Llama 3.1 8B | Fast | Private conversations (runs locally, nothing leaves the homelab) |
| Llama 3.2 3B | Fastest | Quick local tasks where privacy matters |
| Hermes Agent | Slower | Tasks that need tools — web browsing, running code, file work |
Cloud models (Gemini, DeepSeek, Claude) route through OpenRouter. Local models
(Llama) run on Ollama inside the homelab and never send data externally.
The AI remembers things about you across conversations. If you tell it your
preferences, timezone, how you like responses formatted — it stores that and
uses it in future chats.
Memory is per-user and persistent. To see or manage what it remembers, visit
mem0.bvoy.dev.
You can also tell it explicitly: "Remember that I prefer concise answers" or
"Forget what you know about my job."
Any model can search the web. Start your message with something like:
"Search the web for..."
or toggle the Web Search button below the chat input. Searches go through
the homelab's private SearXNG instance — not Google directly.
To generate an image, click the Image icon below the chat input (or ask
directly):
"Generate an image of a misty mountain at sunrise"
Images are generated via Gemini's image model through OpenRouter.
Model: Hermes Agent
Hermes is an agentic model — it can use tools to complete multi-step tasks
rather than just answering questions. Available tools include:
Use Hermes when you need it to do something, not just answer. Example prompts:
"Research the top 5 self-hosted note-taking apps and compare their features
in a table" "Write a Python script to parse this CSV and chart the results"
Hermes runs slower than regular models because it works through multiple steps.
Check hermes.bvoy.dev for session history and tool
logs.