AI infrastructure &
enterprise IT guides
Practical guides on AI infrastructure and enterprise IT: benchmarks, sizing formulas and honest comparisons from our engineering partner’s team.
Four questions decide it: model size, users at once, where the box lives and what follows the pilot. 273 GB/s against 1,597 and 4,800.
Read the article →No VRAM figure, one shared pool: what the memory really holds at BF16, FP8 and FP4, and what a single unit can fine-tune.
Read the article →The vendor publishes no tokens per second for this card. The KV cache arithmetic, the 1,597 GB/s server figure, and how to measure your own.
Read the article →Slot, sense pin, cooling, support list, NVLink and licence: the gates that decide whether 141 GB of HBM3e can go in.
Read the article →A finished backup job is evidence about RPO, not RTO. What a restore leaves behind, and what proven recovery has to show.
Read the article →RAG changes the input, fine-tuning changes the weights. What each one fixes, and what full tuning, LoRA and QLoRA cost in VRAM.
Read the article →How hardened repositories and object lock work, the attacks immutability stops and the ones it does not, and how long to set the window.
Read the article →Standard and Enterprise Plus stop at vSphere 8 U3: what each tier switches on, the per-core rules, the vSAN entitlements and the renewal traps.
Read the article →2.1 TB of HBM3e with NVLink, or 768 GB of GDDR7 on PCIe: memory arithmetic, MLPerf v6.0 results and what each machine is actually for.
Read the article →96, 72, 48, 32 or 24 GB: memory, bandwidth, MIG and vGPU per card, which model fits where, and why CAD does not need the big one.
Read the article →Per GPU, not per VM: how it is counted, what ships with H100 NVL and H200 NVL, when vGPU compute needs it, and what runs without it.
Read the article →Llama 3.3 70B, thirty users, 8k context: the memory arithmetic, why one 96 GB card only just works in FP4, and the server around the card.
Read the article →Where air stops, what rear doors, cold plates and immersion each carry, ASHRAE water classes, flow rates and what a retrofit really involves.
Read the article →Why a 2U takes four 350 W cards but two 600 W ones, where PCIe switches come in, and the one number that stops most orders.
Read the article →How two units link, the measured two-node token rates, image and video generation, fine-tuning, and the honest answer on gaming.
Read the article →96 GB GDDR7 with FP4 or 141 GB HBM3e with NVLink: model footprints, the bandwidth arithmetic behind tokens per second, and the MLPerf numbers.
Read the article →48 GB at 350 W against 96 GB at up to 600 W: memory, MIG, vGPU seats, server fit, and the cases where the older card is the right order.
Read the article →Almost every failure is a retrieval failure. Permissions at retrieval time, chunking that keeps meaning, and the evaluation set nobody builds.
Read the article →The break-even is a token-volume question. The four cost lines nobody counts, and when residency settles it before cost does.
Read the article →Why cost rises steeply as either approaches zero, how to tier systems, and the dependency order that ruins untested recovery plans.
Read the article →Oversized VMs, zombies, forgotten snapshots and stale headroom, and why more vCPU often makes a VM slower.
Read the article →How cores are counted, what the 16-core minimum does to a small cluster, what the 72-core story really was.
Read the article →Rack density numbers, the airflow arithmetic, where air cooling stops working, and the four checks that decide whether your delivery can be switched on.
Read the article →Measured token rates, the 273 GB/s ceiling, batch serving results and two-node pairing: what the spec sheet doesn’t say.
Read the article →Which NX nodes take GPUs, how MIG and vGPU divide them, what works on AHV, and the limits that break multi-GPU plans.
Read the article →One die, three cards: the complete spec table NVIDIA never published in one place, and which version fits which rack.
Read the article →Hardware slices vs time slices, per-card seat counts for RTX PRO and L40S, licences, and where the setup breaks.
Read the article →96 vs 48 GB, FP4 and MIG that Ada never had, and the cases where the previous generation still does the job.
Read the article →Identical compute, all the difference in memory: where the gain reaches 3.4×, and where it is exactly zero.
Read the article →Weights, KV cache, overhead and concurrent users: the working formulas for choosing a GPU for local LLM inference.
Read the article →A sizing question, a configuration, a quote: we reply within one business day