Owner & Builder

Self-Hosted AI Platform

A self-hosted inference platform on a dual-GPU workstation: hardware, model stack, an evaluation loop for new tools, and the working practices behind faster customer prototype turnaround.

  • Ubuntu Server
  • Ollama
  • Qwen 3.8
  • Dual RTX 5060 Ti — 32 GB VRAM
Self-Hosted AI Platform

Overview

This document describes a self-hosted large language model inference platform built on a single dual-GPU workstation, and the working practices that have grown around it. It covers the rationale for running inference locally, the hardware and software stack, the process used to evaluate new models and tools, and the measured effect on delivery work.

Status — in daily use.

Scope — one workstation, one operator, local network only. No external users and no multi-node scaling.

Background — why self-hosted

Managed inference services imposed three recurring costs:

  • Per-token billing. Every experiment carried a marginal cost, which discouraged exploratory work.
  • Vendor-controlled change. Model versions and availability changed on the provider’s schedule rather than the project’s.
  • Data egress. Input data had to leave the local environment to be processed.

Running inference on local hardware removes all three: there is no per-request charge, the model set changes only when the operator changes it, and data does not leave the machine. The trade-off is that capacity planning, model selection, and maintenance become the operator’s responsibility.

Hardware

Component Specification Notes
CPU Intel Core i7-13700K —
Memory 64 GB DDR5 —
GPU 2 × NVIDIA RTX 5060 Ti 32 GB combined VRAM
OS Ubuntu Server —

PCIe layout. The motherboard provides a single x16 slot; the second GPU runs at x4. In practice this has no measurable effect on inference throughput. The reduced lane width only slows the one-time transfer of model weights into VRAM at load; once the weights are resident, both GPUs serve at full speed. Accepting the asymmetric slot layout in exchange for doubled VRAM was a deliberate choice.

Software stack

Ollama serves a working set of open-weight models. The primary model is Qwen 3.8, quantized to fit the 32 GB VRAM budget while leaving headroom for a context window of roughly 100k tokens — large enough to hold a full document in context and still run inference.

Adopting a new model or tool follows a fixed three-step loop:

  1. Pull the model or tool.
  2. Run it against a small, fixed set of realistic tasks that reflect actual work.
  3. Adopt it into a workflow, or discard it.

The platform is not a monolith; it is a set of small, independently revocable decisions.

Division of responsibility

The system performs inference. The operator owns everything around it: tracking new releases, deciding what is worth running, assembling the stack, and connecting it to real work.

The setup itself is straightforward for a single operator. The judgment calls are the substantive part:

  • which model size fits 32 GB of VRAM without unacceptable latency;
  • which tools justify a permanent place in the daily workflow;
  • which experiments are worth pursuing and which are not.

These are senior-level decisions. The local platform makes them cheap to make and cheap to reverse.

Technology evaluation

The platform doubles as an evaluation environment. When a new model, agent pattern, or inference tool is released, it is run directly against real tasks rather than assessed from published benchmarks. A few hours of side-by-side runs on representative work replace roughly a week of reading secondary sources.

This shortens the decision cycle to days: each candidate is adopted, held, or dropped within that window. Failure modes are observed first-hand, on local data, rather than inferred from others’ reports. Because candidates are evaluated as they appear, an adopted tool is already installed and characterised by the time it is needed.

Working practices

Two workflows account for most of the platform’s daily value.

Review support

Code, designs, and documents are passed through the model before human review. It surfaces mechanical issues reliably — missing edge cases, unclear requirements, unstated assumptions — so that human review can concentrate on the judgment calls that require a person.

Customer prototypes

From the point at which requirements are gathered, a working prototype is produced in hours rather than days. Customers see a concrete artifact early, feedback is grounded in something real, and scope estimates are based on a working system rather than projections. Requirements to working prototype within the same working day is the single largest change in how work is scoped and delivered.

Skill development

An unplanned benefit was the effect on the operator’s own skill. Working against a local model taught a repeatable prompting process rather than isolated tricks: iterate against real output, request critique before answers, compare drafts against each other, and keep what holds up. The process transfers to any model and has raised the quality of AI-assisted work generally, not only on this platform.

The same environment compresses how new technology is learned. Dense documentation becomes working examples that can be modified and broken; trade-offs are examined against a running system rather than a written description; and a system is always available to answer the same question asked a different way. Time to a first working prototype dropped from days to hours for most tasks.

Results

  • Customer prototype turnaround: from days, often weeks, to hours, with earlier customer sign-off and scope estimates that hold.
  • Review throughput: mechanical gap-finding that previously required careful human passes is handled before review.
  • Learning time: first working prototype for ordinary problems measured in hours.
  • Cost per experiment: zero marginal cost, so testing is no longer a budget decision.
  • Data handling: input data is processed locally and is not sent to a third party.