Jafar Uruç

Local Inference Models

2026-08-08

One of the most exciting thing in AI is not just that powerful models can solve a fair number of tasks in a single attempt. It is that more of these models can now run locally.

Here are some local models I currently use:

Today, I experimented with MiniMax-H3. I asked Fable, using its xhigh effort setting, to write an optimized inference script. MiniMax-H3 fits easily in the VRAM of my RTX 3090 and takes about 22 minutes to generate a 124-frame video at 864x480 resolution.

I am looking forward to trying Qwen3.8 when it is released.

#llms #local-inference