Local AI inference build
As I was preparing to leave the enterprise Github Copilot unlimited token all-you-can-eat buffet, I started to put a keen eye on token usage. In the early parts of 2026, you couldn’t track this, only the size of the context your session is using. This is related to the amount of input and output tokens but correlation back then was difficult. Then in spring, Github decided it was losing their shirts, pants, socks, etc. because everybody was burning tokens on agentic coding at an eye-watering rate. Then they shut down the unlimited plans and announced they were going to charge via “AI credits”. I just spent a month in the AI credit world and I’ll write about this and other related topics separately.
It has been a dream of mine to build a solar-powered personal AI data lab in my spare bedroom. I started with an nVidia 3060 GPU card with 12GB VRAM for local AI model processing of Photoshop and 3rd party filters. I’ll write about this separately soon. What I’ll say here is that that amount of VRAM and that type of setup (Thunderbold 4 direct connect) is not good for local inference for agentic coding scenarios.
The dream: within 18-24 months, my local inference setup would allow me to have Opus 4.5 level reasoning in a 160K context window.
I quickly abandoned starting out with 1 3090 Founder’s Edition (FE) card (24GB VRAM) and also hosting 1 or 2 cards in a Sonnet eGPU enclosure. The latter has space, heat dissipation and remote access issues. As a result, the host for the two cards is a proper “server” box, in actuality, a very well-built gaming PC.
This resulted in a search for 2 3090 FE cards. Unfortunately, this turned into a discouraging and ultimately failed venture.
The first thing that Gemini showed me is below.

This is a total hallucination (or old training data). Not only is this card not available at this price, it’s also not available at any price from the (semi-)reputable refurb sites. And below that level is the seedy underbelly of the Internet where sketchy is everyone’s middle name. Burned-out Bitcoin mining card, anyone? (um, no).
The 3090 FE is an older card, but it is the sweet spot for people trying to do what I am wanting to do. It is theoretically available on enterprise teardowns, but only to people who are properly connected. I went to UNIXSurplus and PC Server & Parts who are “industrial recyclers or data center decommission remarketers.” But I sent both of them specific requests via email and they just ignored me.
I went weeks failing to secure even one of the two desired 3090 FE cards and was lamenting that this project would fail before it even got started or cost several times more money than originally budgeted (e.g. buy new 5090 cards). I had this fantasy that I could set this up for $4,000 to $5,000.
Somehow, in one of the many LLM conversations I had doing research, I stumbled across CustomLuxPC’s, a custom builder in New Jersey. And, under the “machine learning” section of their website, lo and behold I saw this:

These are not (necessarily) “Founders Edition” GPU’s, which are highly sought after due to superior cooling configuration and capability, but it’s a machine learning rig with dual 3090’s in it. The web page has a “text us” SMS link and this link is for real. You can text Pat and he responds very quickly. It became apparent right away that not only is Pat knowledgeable, but this is also not some fly-by-night sketchy operation.
I won’t recount our conversation, but I had a lot of questions and a few requirements and Pat answered them all, knowledgeably and cheerfully. I did detour into “would one A6000 card (48GB VRM) be better?” but ultimately decided the added expense of that would be left for another day.
You can read about the configuration on the link above, I upgraded to FE cards and 2×64 DDR5 RAM (instead of 4×32) so I can upgrade to 256GB RAM later.
The build is done. This is my new AI data center PC. It shipped today. 🎉❤️

The “status” mail I received said:
“The PC is ready for shipping. It passed all GPU, CPU, and memory tests with flying colors and the system stays cool under load. The top GPU stays around 65-70 C while the bottom GPU is 58-65 C. The CPU maxes out at 75 C under full load.”
I haven’t owned a desktop format factor PC in literally decades. I’ve been using laptops and high-powered cloud VM’s (the latter being my “dev boxes”). You should be able to tell quickly that I am not a PC gamer, either.
In future posts, I’ll delve into initial setup and how I am configuring this using (I think?) Ollama, 4-bit quantized models, TailScale/WireGuard and yet another VLAN on my Ubiquiti UDM7 router. The power for this will be supplied by 24 solar panels and the UPS is a whole-house battery (basically I just plug this setup directly into the wall). Instead of putting all the excess solar power into the grid, I plan on using some of that excess to power this setup.
I can’t say enough good things about the experience I had with CustomLuxPC’s. I have every confidence in the build and, more importantly, that there will be support should I need it. Pat turned a dream into reality.