LIBRISTO
LIBROAMANTO
obvezno
Postanite del skupnosti ljubiteljev knjig z vsega sveta in uživajte v številnih ugodnostih. Ustvarite brezplačen račun
0
Brezplačna dostava Zásilkovna nad 69.99 €
Zbirna točka GLS 4.49 € Zbirna točka DPD 2.99 € Kurirska služba GLS 5.49 € Kurir DPD 3.49 € Kurirska služba Express One 3.49 € Zbirno mesto Express One 3.49 € Zbirno mesto Pošte Slovenije 3.49 € Dostava preko Pošte Slovenije 3.49 €

Brezplačna dostava za naročila nad 69,99 € na prevzemna mesta DPD in Express One.

LLM Inference Engineering

From KV Cache to Production Serving: A Practical Guide, Basics to Advanced

Jezik AngleščinaAngleščina
Knjiga Mehka
Knjiga LLM Inference Engineering Shashank Nainwal
Koda Libristo: 54074733
Založba Independently published, oktober 2026
Make your large language models faster, cheaper, and ready for production.Training a model happens o... Celoten opis
? points 45 b Novo Novo
18.64 €
Na zalogi pri dobavitelju Odposlali bomo v 10-16 dneh

Do 30 dni za vračilo

Make your large language models faster, cheaper, and ready for production.

Training a model happens once. Serving it happens every time someone uses your product, and that is where most AI budgets go. This book explains, step by step, what really happens when an LLM generates text and how engineers make it fast and affordable at scale.

Starting from a single question, "what happens when a model produces one token?", you will build a complete mental model of modern inference, from first principles to production deployment.

What you will learn


  • Why decoding is memory-bound, and how to predict speed and cost with simple napkin math

  • The KV cache, PagedAttention, FlashAttention, and GQA/MQA/MLA explained in plain language

  • Continuous batching, chunked prefill, speculative decoding, and prompt caching

  • Quantization (FP8, INT4, GPTQ, AWQ), distillation, Mixture of Experts, and model routing

  • Tensor, pipeline, and expert parallelism, plus disaggregated serving and long context

  • How modern serving engines and the GGUF format work, and when to choose each

  • How GPUs and TPUs work, and how to compare accelerators for inference

  • Benchmarking, observability, cost modeling, and capacity planning



Inside the book


  • 27 chapters across 8 parts, from basics to advanced

  • 29 original diagrams

  • Key takeaways and quiz questions in every chapter, with an answer key

  • 3 hands-on labs you can run on a laptop or a single GPU

  • 25 interview questions with model answers

  • A formula cheat sheet, glossary, and curated list of foundational papers



Who this book is for

Software and ML engineers, platform and MLOps engineers, solution architects, technical product managers, and anyone preparing for AI infrastructure interviews. You need basic Python and a rough idea of what a neural network is. No CUDA or GPU required.

Written by an AI architect and former Amazon Web Services engineer who has built LLM applications serving millions of customers.

Igralka & Poliglotka
EWA KASP za
Predvajaj video
Ewa Kasp
Libristo ima največjo izbiro tujejezične literature. Zato svoje knjige kupujem tukaj.

O knjigi

Polni naslov LLM Inference Engineering
Jezik Angleščina
Vezava Knjiga - Mehka
Datum izida 2026
Število strani 130
EAN 9798178522462
Koda Libristo 54074733
Teža 240
Mere 178 x 254 x 7
Podarite to knjigo še danes
To je povsem preprosto
1 Dodajte knjigo v košarico in izberite dostavo kot darilo 2 V zameno vam bomo poslali kupon 3 Knjiga bo dostavljena na naslov obdarovanca

Prijava

Prijavite se v svoj račun. Še nimate računa Libristo? Ustvarite ga zdaj!

 
obvezno
obvezno

Še nimate računa? Izkoristite prednosti računa Libristo!

Z računom Libristo boste imeli vedno vse pod nadzorom.

Ustvarite račun Libristo
Knjižni svetovalec Libroamiko
Pozdravljeni, sem Libroamiko, vam lahko pomagam?