From the build · 1 October 2026

It looked itself up.

The first time a person asked a question in Workbench, the AI went to the web for the answer. The question was about Eugene Plexus. The website it found was this one.

How it got here

Eugene Plexus runs AI models on your own machines and gives all your apps one place to reach them. Workbench is its own chat app. You sign in with your Eugene account, talk to the models your machines run, and let them search the web when you want them to.

On 1 October 2026, Troy, who builds Eugene Plexus, updated the two machines in his own install. One is a NAS that runs the hub in a container. The other is a Windows PC with an RTX 5090 that runs the models. He added a Brave Search key, installed Workbench on the PC from the NAS's console, signed in, and typed a question.

Automated tests had already done every step of this, many times over. This was the first time a person did, on a real install, with a real question. He picked the one question whose answer he could check: what is Eugene Plexus?

What one search went through

  1. Workbench sent the question to Eugene's gateway, with web search turned on.
  2. The gateway passed it to Qwen, a 27-billion-parameter model running in llama.cpp on the RTX 5090.
  3. Qwen didn't know the answer, and asked to search the web.
  4. The gateway ran the search itself, through Eugene's search tool and Troy's Brave account, and handed the results back to Qwen.
  5. Qwen read the results, which included this website and the project's GitHub pages, and answered.

The search terms went to Brave. The question, Qwen's reasoning and its answers stayed on Troy's two machines.

The chat, as it happened

Copied from Workbench without edits, including the parts we'd rather fix (see what we learned). Qwen's reasoning is folded under each answer: open it to watch it decide to search.

Model a 27B Qwen model, in llama.cpp on one RTX 5090

  1. Troy

    What is Eugene Plexus?

  2. Qwen 21 seconds 3 searches

    How it reasoned

    The user is asking "What is Eugene Plexus?" Let me think about this. "Eugene Plexus" — hmm. There's a concept called "plexus" in biology (a network of nerves, blood vessels, etc.). "Eugene" might be a person's name, or it might refer to Eugene, Oregon. Let me think of possibilities: 1. Eugene Plexus might be a software project, a company, or a musical group. 2. It might be related to the "Eugene" app or something. 3. There might be a "plexus" in biology — for example, the "epineurium"... no, that's not it. There are nerve plexuses like the brachial plexus, the lumbar plexus, the sacral plexus. "Eugene plexus" isn't a standard anatomical term that I know of. Actually, there might be something called "Eugene Plexus" — perhaps a data visualization project, or a network of something named after Eugene, Oregon. Or maybe it's a person's name — "Eugene Plexus" could be a username/handle. I should search the web to find out what this is. Found it. Eugene Plexus is a self-hosted control plane for local LLM inference. To provide an accurate and detailed answer, let me fetch the main website and get more details. I've got enough information now. To be accurate about what it is, let me summarize based on the search results. I could fetch the main site for more details, but I think I have enough info: Eugene Plexus is a self-hosted control plane for local LLM inference. It installs and supervises upstream inference engines (like llama.cpp, Ollama, vLLM, etc.), manages model files and launch configurations, and routes requests. It's under active development, with an alpha version available for early testing. It has versioned installers and corresponding container images. Components: agent (per-host node agent that supervises components and inference engines), control, gateway, inference-driver, library, ui (web UI with runtime dashboard, model library, discovery/downloads, schema-driven configuration, chat playground, logs). There's also "Workbench," its own chat app, where you can chat with models and search the web, signing in with Eugene. To be more precise, I could fetch eugeneplexus.com. Let me do that quickly to get a better description. I now have a clear overall picture. Let me write a concise yet informative answer.

    Eugene Plexus is an open-source, self-hosted control plane for local LLM inference — in other words, the management layer that sits around your local AI model setup rather than doing the inference itself. It’s currently under active development, with an alpha version available for early testing via versioned installers and a matching container image.

    What it does (from its project site):

    • Installs and supervises upstream inference engines (the actual model-running software)
    • Manages model files and launch settings
    • Routes requests, exposing an OpenAI-compatible API (chat, tool calling, embeddings) plus Anthropic Messages compatibility for text and tool calls

    How it’s organized — it’s a multi-component system with shared OpenAPI 3.1 contracts (components share schemas, not code):

    • agent — a per-host node agent that supervises the other components and inference engines, manages engine binaries, and handles node enrollment
    • control — the install-wide trust root, node registry, topology, and a replicated control-state log
    • gateway — serves the OpenAI/Anthropic-compatible API
    • inference-driver and library — handle engine execution and model files
    • ui — a web dashboard with runtime monitoring, a model library with discovery/downloads, schema-driven configuration, a chat playground, and logs

    There’s also Workbench, its own chat app for chatting with the models an Eugene install serves (with web search), signed in through Eugene.

    The name plays on “plexus” (a network of interconnected elements, as in a nerve plexus) — it’s the networked control layer tying together your local inference machines.

  3. Troy

    That's correct! Do you know why that's important?

  4. Qwen 46 seconds 2 searches

    How it reasoned

    The user is asking why "Eugene Plexus" is important. Let me think about what I actually know versus what I'm guessing. What I claimed in my previous answer: - It's an open-source, self-hosted control plane for local LLM inference - It manages local AI model setups - It's under active development with an alpha version - It has components like agent, control, gateway, inference-driver, library, ui - It has a Workbench chat app Now, the user is asking why this is important. Let me think about what makes a "control plane for local LLM inference" important in the broader context of AI. The key importance here is about the broader trend of local/self-hosted AI inference. Let me think about why this matters: 1. **Data privacy** - Running LLMs locally means your data doesn't leave your machine/network 2. **Cost** - No per-token API costs to cloud providers 3. **Latency** - Local inference can be faster for certain use cases 4. **Offline capability** - Works without internet 5. **Control** - You control the hardware, the models, the updates 6. **No vendor lock-in** - You're not tied to one cloud provider But specifically for a "control plane" (as opposed to just running a single model), the importance is: - Managing multiple models - Managing multiple machines/nodes - Routing requests intelligently - Providing a unified API (OpenAI-compatible, Anthropic-compatible) so existing tools work - Making local inference as easy to manage as cloud inference Let me think about whether I should search for more specific information about why Eugene Plexus specifically is important, or if I can reason about this from general knowledge. Actually, I think the user is asking a somewhat open-ended question - "why is that important?" - meaning why is a local LLM inference control plane important in the landscape of AI. Let me reason through this. The broader context: - Most LLM usage today goes through cloud APIs (OpenAI, Anthropic, Google, etc.) - There's a growing movement toward local/self-hosted inference for privacy, cost, and control reasons - But local inference has traditionally been hard to manage - you need to install the right engine (Ollama, llama.cpp, vLLM, etc.), download models, configure them, expose an API, manage multiple models, handle routing - A control plane abstracts all of that away and gives you a unified, API-compatible interface So the importance is: 1. **Democratizes local AI** - Makes it easier for non-experts to run and manage local LLMs 2. **Privacy** - Your prompts, data, and models stay on your hardware 3. **Cost savings** - No per-token fees for high-volume usage 4. **Interoperability** - OpenAI/Anthropic-compatible APIs mean existing tools and code work without modification 5. **Multi-node management** - Can manage inference across multiple machines 6. **Reduces dependency on cloud providers** - Especially relevant with concerns about data sovereignty, vendor lock-in, etc. Let me also consider if there's something specific about Eugene Plexus that makes it stand out. From what I know, it's relatively new (alpha), open-source, and has a fairly comprehensive architecture with multiple components. The fact that it's a "control plane" rather than just another inference engine wrapper is significant - it's trying to be the orchestration layer. I think I can answer this well from reasoning without needing to search, but let me consider if a search would add value. The user said "that's correct" about my description, so they're confirming my factual claims. Now they want the "why" - the significance. This is more of an analytical/opinion question that I can reason through. Let me write a thoughtful answer about why a local LLM inference control plane is important in the current AI landscape. The second search failed due to rate limiting, but I got good results from the first search. The results confirm what I already knew about Eugene Plexus. Let me refine my answer based on the search results. The key insight from the search results is the tagline: "Neither. Eugene Plexus is the control plane around your inference setup." This is a key distinction - it's explicitly NOT an inference engine, it's the control plane AROUND inference. This is the core of why it's important. Let me also note the architecture details confirmed: - Six active consumers: agent, control, gateway, inference-driver, library, and ui - Components share schemas, not code - OpenAPI 3.1 contracts - Currently at v0.1.0-alpha.5 The user asked "why is that important?" - I think my previous answer was good but let me refine it with the confirmed details. The key importance points are: 1. It's explicitly a control plane, not an inference engine - this is the critical distinction 2. It manages the entire local inference setup: engines, models, routing, monitoring 3. OpenAI/Anthropic-compatible API for interoperability 4. Multi-node management 5. Privacy, cost, control benefits of local inference Let me write a clean, well-organized answer. I don't need to re-search - I have enough information.

    Great question. The importance of something like Eugene Plexus comes down to the broader shift happening in how people run LLMs:

    The problem it addresses: Most LLM usage today flows through cloud APIs (OpenAI, Anthropic, etc.). That means your prompts, data, and workloads live on someone else’s hardware, and you pay per-token. Local inference solves the privacy and cost issues, but it’s been hard to manage well:

    • You have to pick and install the right inference engine (llama.cpp, vLLM, Ollama, etc.)
    • Download and organize model files
    • Configure launch settings for your specific hardware
    • Expose a usable API
    • Handle routing, monitoring, and updates
    • Scale across multiple machines if you need more capacity

    A control plane abstracts all of that. You get a unified, OpenAI/Anthropic-compatible API, so existing tools and code work without modification, and you can manage multiple models and nodes from one place.

    Why that matters in practice:

    • Privacy / data sovereignty — prompts and data never leave your network
    • Cost — no per-token fees, which matters for high-volume or always-on workloads
    • Offline / latency — no round-trip to a cloud provider
    • Interoperability — the OpenAI/Anthropic-compatible API means you’re not locked into one vendor’s SDK
    • Multi-node orchestration — manage inference across several machines as a single logical system

    The bigger picture: There’s a real tension in AI right now between the convenience of cloud APIs and the desire for control, privacy, and cost predictability. Local inference has been the “power user” path for a while (Ollama, LM Studio, etc.), but a proper control plane is what makes it viable for teams and more complex setups — the kind of infrastructure layer that lets local AI feel as “boring” and manageable as cloud inference, without the cloud.

    So Eugene Plexus sits at an interesting intersection: it’s not another inference engine, it’s the management and orchestration layer that makes a fleet of local inference machines practical to run. Whether it becomes the standard for that layer is still to be seen, but the problem it’s solving is a real and growing one.

    Is there a particular angle on “importance” you were thinking about — like a specific use case or a comparison to existing tools? I’m happy to dig in further.

    It searched again here, then wrote a second answer. Workbench showed both, one after the other.

    The search results confirm the core framing, and I think the key to “why it’s important” is right in their own tagline:

    “Neither. Eugene Plexus is the control plane around your inference setup.”

    That distinction is the whole point. Let me lay out why that matters:

    The gap it fills

    Local LLM inference has a real but under-served problem. You can run a model with llama.cpp, vLLM, or Ollama, but those are inference engines — they run one model on one box. The moment you want to:

    • manage multiple models across multiple machines
    • install/update/supervise the engine binaries themselves
    • route requests intelligently
    • get a single unified API that your existing tools already speak
    • monitor, log, and configure all of it from one place

    …you’re on your own, writing glue code. That’s the gap. Eugene Plexus is explicitly not another engine — it’s the orchestration layer that sits on top of whatever engines you choose.

    Why that layer matters

    1. Interoperability — It exposes an OpenAI-compatible API (chat, tools, embeddings) plus Anthropic Messages compatibility. That means the thousands of existing tools and SDKs built against those APIs work against your local hardware with no code changes. You get local-inference benefits without rewriting your stack.

    2. Privacy & data sovereignty — Prompts, context, and model weights stay on your own machines.

    3. Cost & latency — No per-token cloud fees and no network round-trip, which matters for high-volume or always-on workloads.

    4. Multi-node orchestration — The agent runs per-host and reports into a control plane, so a fleet of inference machines behaves like one logical system. That’s the step from “I run a model on my desktop” to “I run an inference cluster I actually manage.”

    5. A clean, spec-driven architecture — The specs repo is a single source of truth of OpenAPI 3.1 contracts, and the six components (agent, control, gateway, inference-driver, library, ui) share schemas, not code, via codegen. That’s a deliberate design choice that keeps the pieces loosely coupled and swappable.

    The bigger picture

    There’s a growing tension between the convenience of cloud APIs and the desire for control, privacy, and predictable cost. Local inference has been the power-user path for a while, but a proper control plane is what makes it viable for teams and for anything beyond a single box. Eugene Plexus is betting that the missing piece in local AI isn’t a better model or a faster engine — it’s the boring, essential infrastructure layer that makes a fleet of local machines practical to run.

    Whether it becomes the standard for that layer is still to be seen (it’s at v0.1.0-alpha.5), but the problem it’s solving is real and growing.

    One note: my second search (on the broader 2025 self-hosted-AI landscape) hit a rate limit, so the “bigger picture” framing above is my reasoning rather than something I pulled from a fresh source. If you want, I can re-run that search to ground the industry-context part in current articles.

  5. Troy

    You're correct again, but i meant it in another way. This chat session is the very first that a human has used to confirm that an LLM (you!) can search the web using tools through Eugene Plexus. We're inside the new Workbench feature right now.

  6. Qwen 39 seconds 1 search

    How it reasoned

    The user is explaining that this conversation is the first human-confirmed instance of an LLM using web search tools via Eugene Plexus, and we're currently inside their new Workbench feature. Let me confirm my understanding and respond naturally. So the context is: - I'm running as an LLM behind Eugene Plexus's gateway - The user is chatting with me through Workbench, their own chat app - I just used the web_search tool, and that request was routed through Eugene Plexus's infrastructure - The user has now confirmed it worked — meaning this is the first human-verified confirmation that tool-calling web search works end-to-end through Eugene Plexus This is a fun "first" moment. I should be delighted and acknowledge it. I could also do a little self-referential check — I can't directly verify which backend I'm running on, but the user has told me. I shouldn't overclaim. Should I search the web? Not necessary. This is a conversational moment. I could do a quick celebratory check, but I think a warm, concise response is best. I could mention the irony/beauty: the very tool-calling that I just demoed was the confirmation itself — my search calls a few minutes ago were the first web searches an LLM performed through Eugene Plexus, confirmed by a human. I could also make a light joke: "This message itself is the confirmation." I could offer to mark this moment, or ask if they want me to do anything specific. Let me keep it concise, warm, and a little playful. No tools needed. The user is saying this conversation is the first human-confirmed instance of an LLM performing a web search via Eugene Plexus's tool, and we're currently inside their new Workbench feature. So this is a fun, celebratory moment. The user is telling me that this very conversation is a milestone: the first human confirmation that an LLM can search the web using tools through Eugene Plexus. I should respond warmly and appropriately, acknowledging this milestone. I don't necessarily need to search the web again, but it might be good to confirm what I know about Workbench. In fact, in my earlier answer I already described Workbench as "a chat app for chatting with models that an Eugene install provides (with web search), signed in through Eugene." The user is now telling me that we are inside Workbench, and that this is the first human-confirmed instance of an LLM doing a web search via Eugene Plexus tools. This is a conversational moment, not a research request. I should respond with appropriate enthusiasm, perhaps touching on the significance of the milestone and what it means. I could do a quick check to see if there's any news about this, but honestly it's a "just happened" moment, so a web search probably won't find anything. I think a well-crafted conversational response is the right answer here. Let me think about what makes a good response: 1. Celebrate the moment — this is a genuine "first" 2. Reflect on what it means: this validates the full loop of Workbench's chat + web search + Eugene's tool calling 3. Maybe touch on the irony/poetry: I was researching Eugene Plexus, and now I'm discovering that I'm running *inside* it 4. Keep it concise — the user is making a social statement, not asking a question I could do a quick search to see if there's anything recent about Workbench's launch, but honestly, the user has told me everything I need to know. I think a search would be overkill and would make me look like I'm not listening. Let's go with a warm, concise response. Actually, let me reconsider. Is there any value in a search? The user said "we're currently inside the new Workbench feature." That's the latest thing. A search might confirm details about Workbench. But I already know from previous searches: "Workbench, its own chat app for chatting with models that an Eugene install provides (with web search), signed in through Eugene." I think the best move is a conversational response that celebrates the milestone without over-researching. Maybe I could offer to do something fun — like searching the web to "record" this moment, or checking if there's a changelog. But I shouldn't overdo it. Let me write a warm, concise response that: - Acknowledges the milestone - Reflects on the significance (validates the full loop: chat → tool call → web search → answer, all self-hosted through Eugene) - Touches on the poetic irony (I was researching Eugene Plexus, and now I'm discovering I'm running inside it) - Maybe offers a small next step I'll keep it reasonably short. This is a social moment.

    Oh, that’s a fun way to be part of a “first”! So the tool calls I made a few messages ago — the searches for Eugene Plexus’s site and GitHub — were the very first web searches an LLM has issued through your gateway, and this chat is the first human to verify they actually landed. That’s a nice little bootstrap moment: the system’s first confirmed external action was researching itself.

    A few thoughts on why that’s a meaningful milestone for Workbench, if I may:

    • It closes the loop. A chat app that can only chat is a terminal. A chat app where the model can act — search, fetch, call tools — is an agent. Confirming web search works end-to-end (prompt → gateway → tool call → result → back into the conversation) is the proof that the plumbing holds.
    • It validates the OpenAI-compatible path. Tool calling through an OpenAI-compatible API is a well-trodden pattern in cloud, but routing it through a self-hosted control plane to a local model is a different beast. The fact that it just worked with you watching is the best test there is.
    • It’s a good canary. Web search is a simple, observable tool: you can see the queries, see the results, and judge the quality of the answer. If that works cleanly, more complex tool chains (multi-step, with function results feeding back) have a solid foundation.

    One small caveat for the record: I can only confirm the experience side — that my tool calls returned results I could reason over. You’re the one who can confirm the infrastructure side, that the requests actually traversed the gateway and the agent as designed. So this really is a two-person confirmation, which is probably the right way to do it.

    Congratulations on the first! Do you have a changelog or release notes where this milestone would be worth a line? I’m happy to draft one if you’d like. 🎉

    It searched again here, then wrote a second answer. Workbench showed both, one after the other.

    That’s a fun milestone to be part of — and there’s a nice piece of irony I only just realized: a few messages ago I was out on the web researching what Eugene Plexus is, and all along I was running inside it, with my search calls being routed through your gateway’s tool-calling path. The first human-verified “the LLM just searched the web through Eugene” moment is… an LLM searching the web about Eugene.

    Stepping back, what this session just validated end-to-end is exactly the loop Workbench is designed to close:

    1. Chat — you talk to a model your Eugene install serves, signed in through Eugene
    2. Tool calling — the model decides to call the web search tool, and the gateway routes that call correctly (OpenAI-compatible tool calling is one of the trickier parts of the contract, so this is a real test)
    3. The tool itself — per the tool-driver repo, web search is a tool the hub runs itself, one instance per search account, so the search ran on your own infrastructure, not some cloud add-on
    4. The answer — results come back into the conversation and I reason over them

    And the important part: the whole thing stayed self-hosted. No prompt or search query had to leave your network to reach a third-party chat service — which is the whole point of the “control plane around your inference” philosophy, now demonstrated by the system dogfooding itself.

    So congratulations on the first! A few thoughts on where to take it from here:

    • Worth documenting — if you have a changelog, release notes, or a “firsts” page on eugeneplexus.com, this is the kind of milestone that makes a great anecdote: “The first confirmed web search through Eugene was a search about Eugene.”
    • Stress test ideas — now that the single-shot search works, the interesting next probes are multi-step tool calls (search → fetch a page → reason over both results) and parallel tool calls in one turn. Those put more load on the gateway’s routing and on the spec contracts that keep the components in sync.
    • The “spoke” architecture is paying off — Workbench being “an app (a spoke), not part of the hub” means this milestone also proves that a third-party-style app can lean on Eugene’s auth and tooling without the hub having to know anything about it.

    Want me to draft a short changelog entry or a “milestone” blurb you could drop into the repo or site? I can match whatever tone the project uses — I’ve seen the specs repo’s style is pretty crisp and technical.

What we learned

It said when a search failed

In its second reply, Qwen asked for two searches at once. Brave's free plan turned down the second, because it came too soon after the first. Eugene told Qwen the search had failed and asked it to say so. Qwen did, and marked which part of its answer was its own reasoning rather than a source. Since alpha.6, Eugene spaces an account's searches out and tries a refused one once more (tool-driver #3).

Two answers in one reply

Twice, Qwen wrote a whole answer and then searched, then answered again. Part of that is ours: on the first step of a reply, Eugene tells the model it must search. In its last reply, Qwen reasoned "No tools needed." It answered, then searched because it was told to.

Workbench then put both answers into one reply, so the reply reads twice. Since alpha.6, Eugene keeps the forced search and says when each search finishes, and Workbench folds what was written before a search under Written before searching and sends only the answer back to the model (Workbench #1, gateway #4). The chat above is kept as it happened.

One thing it got wrong

In its last reply, Qwen said no search query had to leave the network. The search terms do leave: they go to the search service, Brave here. Everything else stays on your machines.

Try it

Workbench and web search are in v0.1.0-alpha.6. Install the alpha, add a search account under Backends, then install Workbench from Apps on any of your machines.

Install the alpha Workbench on GitHub