Sponsored Content

DEV Community

#vllm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

1
Comments
11 min read
vLLM v0.28.0: the breaking change small GPU users must read

vLLM v0.28.0: the breaking change small GPU users must read

Comments
5 min read
Installing Rust for vLLM on Graviton: a G5g walk-through 🦀

Installing Rust for vLLM on Graviton: a G5g walk-through 🦀

Comments
10 min read
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Comments
9 min read
DGX Spark (GB10) bare-metal vLLM: the install that works, two landmines, measured timings

DGX Spark (GB10) bare-metal vLLM: the install that works, two landmines, measured timings

Comments
2 min read
Ollama to vLLM: When to Migrate Your Local LLM Server

Ollama to vLLM: When to Migrate Your Local LLM Server

Comments
15 min read
The unofficial TPU migration guide: Cloud TPU API to Compute Engine

The unofficial TPU migration guide: Cloud TPU API to Compute Engine

6
Comments 2
17 min read
Self-Hosted Gemma 4 on TPU v6e: Deployment & SRE with Antigravity

Self-Hosted Gemma 4 on TPU v6e: Deployment & SRE with Antigravity

Comments
8 min read
What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

Comments
5 min read
Serving Gemma 4 E2B on a TPU v6e-1: what Trillium buys, and what it doesn't

Serving Gemma 4 E2B on a TPU v6e-1: what Trillium buys, and what it doesn't

2
Comments
20 min read
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

13
Comments
9 min read
Serving Gemma4 with Rust on vLLM 🦀

Serving Gemma4 with Rust on vLLM 🦀

10
Comments
10 min read
Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1

Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1

16
Comments 1
21 min read
Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)

Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)

Comments
3 min read
vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM

vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM

Comments
6 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.