Tools
Private browser test

Can my computer run local AI?

Run a quick private test of your memory, processor, and graphics

About 3 seconds · runs only in this browser

MemoryProcessorGraphics

The speed test and its measurement stay in this browser. Results are a practical estimate, not a guarantee.

How local AI works

What happens when AI runs on your computer?

A model file lives on your computer. When you ask something, your hardware runs that model and creates the answer there.

You ask a question

Your prompt goes straight into the local app on this computer

Your computer runs the model

Its CPU, GPU, and memory generate the answer one piece at a time

No internet needed

Your prompt and reply stay here with no cloud, account, or connection required

Inside Taby

A small local brain, tuned for everyday help

Taby runs fine-tuned Gemma-family models in 2B, 4B, and 12B sizes. Smaller models keep everyday tasks, notes, planning, and assistant replies quick, while the 12B option gives stronger answers when more memory is available. It is built for efficient local use, with cloud help still optional.

Quick questions

Can my computer run local AI?

Most modern computers can run a small local AI model. The best model size depends mainly on available memory, CPU speed, and GPU or unified-memory performance.

How much RAM do I need for local AI?

8 GB can run small quantized models. With enough free memory and a short or moderate chat context, 16 GB can usually try an 8B model in 4-bit form, 24 GB can try many 12B models, and 32 GB can try some 27B models. Larger options may run much slower when they do not fit in GPU memory.

Do local AI models use RAM or VRAM?

They can use either or both. System RAM decides what the CPU or a hybrid setup can hold. Dedicated GPU memory, called VRAM, decides how much can stay on the graphics card and run faster. If a model fits in RAM but not VRAM, local runtimes can split work between the CPU and GPU. Apple silicon uses unified memory shared by both.

Can a Mac run local AI?

Yes. Apple silicon Macs are particularly good at local AI because the CPU, GPU, and memory work closely together. Model size still depends on how much unified memory your Mac has.

Does this test upload my computer information?

The WebGPU or CPU speed test runs inside your browser, and its measurement is not uploaded. The website may still collect normal page analytics, and it does not download a full AI model.

Does this local AI checker work in Firefox?

Yes. Firefox can run the private speed test. It uses WebGPU when Firefox and the operating system support it, then falls back to a local CPU test when they do not. Firefox does not share system RAM with websites, so you need to choose your RAM manually to finish the model recommendation.

Can this browser check my VRAM automatically?

Not reliably. Browsers can reveal a broad graphics name and WebGPU support, but usually hide the exact dedicated GPU memory. On Windows, open Task Manager with Ctrl + Shift + Esc, choose Performance, select the dedicated GPU, and read Dedicated GPU memory. You can then enter that amount in the checker.

What can the browser tell us?

Usually it can see the operating system, CPU threads, a broad graphics name, and WebGPU support. Chromium browsers may also share an approximate memory amount, while Firefox does not. Exact VRAM, the precise Apple chip, and total unified memory may be hidden. Apple silicon shares memory between its CPU and GPU, so there is no separate VRAM number to read.

What else affects local AI performance?

Model size and quantization, context length, free memory while other apps are open, available disk space, and the native AI backend all matter. This browser test cannot verify Metal, CUDA, DirectML, native drivers, or your remaining disk space, so its result is a practical estimate rather than a guarantee.

How is the result calculated?

The test starts with system RAM or unified memory because the model, chat context, and local runtime all need room to load. Dedicated VRAM helps estimate how much can stay on the GPU, and the private WebGPU or CPU benchmark helps describe likely response speed. It assumes one 4-bit GGUF model and about a 4K context.

See how Taby uses local AI