Thomas LynchFounder & CEO

The Fat Long Tail of Neoclouds

A neocloud provider is, at its simplest, three things:

  1. A pile of GPUs
  2. Software that operates them as a unified system
  3. An API through which customers use them

The market has believed that the pile is the business. It is not. The pile is a commodity that Nvidia manufactured. The software and the API are the business. Actual provides those pieces. The customers provide the pile.

One install, one fleet

Run a one-line install on each machine and Actual adds it to a unified, heterogeneous fleet. Inference, monitoring, and access across that fleet are all managed for you through a single load-balanced endpoint: api.actual.inc. Enterprises can run the entire platform inside their own environment.

The software is the hard part

vLLM is only meant for datacenter environments. It is a 10GB torch container, and it is more complicated than Kubernetes. Llama.cpp will run on anything, but it cannot serve real volume it lacks batching and robust KV management.

Actual's engine is a small Rust binary built between those two, made to serve production inference on Macs, graphics cards, or DGX Station and Spark.

Who wants their own neocloud?

The fat long tail.

There are maybe 50 companies on earth that can fill a GPU datacenter. There are maybe 500,000 that can fill a rack. An enterprise might use 10-20 DGX Stations. One pro user can bring the same compute in DGX Sparks. At that scale, the customers look the same.

Arming everyone else

The neoclouds buy the pile and rent it back to you. Actual arms everyone else.

Shopify proved this shape with merchants. The long tail wasn't thin, it was fat. Sellers who looked small individually would together dwarf the flagships.

The customer brings the GPU pile. Actual provides the software and API that makes the pile useful.

The product is in beta try it out today: https://actual.inc