Skip to product information
LLM Engineering
LLM Engineering
Description
Book Introduction
A Practical Guide to All Things LLM Engineering

"LLM Engineering" provides a detailed guide to the engineering methods required to develop and deploy production-grade LLM applications. It systematically explores the LLM lifecycle, covering key concepts and practical techniques from data engineering to supervised learning fine-tuning, model evaluation, inference optimization, and RAG pipeline development.
In this course, you will implement an AI that mimics an individual's writing style and personality through a real-world project called 'LLM Twin', and gain in-depth knowledge of practical LLM engineering know-how such as data collection, preprocessing, and model fine-tuning.
By following the practical roadmap presented in this book, you will learn the entire process, from data collection to model optimization, step by step, and take your LLM engineering skills to the next level.
  • You can preview some of the book's contents.
    Preview

index
CHAPTER 1 Understanding LLM Twin Concepts and Architecture

_1.1 LLM Twin Concept
_1.2 LLM Twin Product Planning
_1.3 Development of ML systems based on feature, learning, and inference pipelines
_1.4 System Architecture Design for LLM Twin
_summation
_References

CHAPTER 2 TOOLS AND INSTALLATION

_2.1 Python Ecosystem and Project Installation
_2.2 MLOps and LLMOps Tools
_2.3 Database for storing unstructured data and vector data
_2.4 Preparing to Use AWS
_summation
_References

CHAPTER 3 DATA ENGINEERING

_3.1 Designing the LLM Twin Data Collection Pipeline
_3.2 Implementing the LLM Twin Data Collection Pipeline
_3.3 Collecting raw data into a data warehouse
_summation
_References

CHAPTER 4 RAG CHARACTERISTICS PIPELINE

_4.1 Understanding RAG
_4.2 Advanced RAG Overview
_4.3 LLM Twin's RAG Characteristic Pipeline Architecture
_4.4 Implementing the RAG feature pipeline for LLM Twin
_summation
_References

CHAPTER 5 Fine-Tuning Supervised Learning

_5.1 Creating a Directive Dataset
_5.2 Creating your own directive dataset
_5.3 SFT technique
_5.4 Practical Fine Tuning
_summation
_References

CHAPTER 6 Fine-tuning using preference sorting

_6.1 Understanding the Preference Dataset
_6.2 Creating a preference dataset
_6.3 Preference sorting
_6.4 DPO Implementation
_summation
_References

CHAPTER 7 LLM Evaluation

_7.1 Model Evaluation
_7.2 RAG Evaluation
_7.3 TwinLlama-3.1-8B Evaluation
_summation
_References

CHAPTER 8 Inference Optimization

_8.1 Model Optimization Strategy
_8.2 Model Parallel Processing
_8.3 Model Quantization
_summation
_References

CHAPTER 9 RAG Inference Pipeline

_9.1 Understanding LLM Twin's RAG Inference Pipeline
_9.2 Exploring Advanced RAG Techniques in LLM Twin
_9.3 Implementing the RAG Inference Pipeline for LLM Twin
_summation
_References

CHAPTER 10 Deploying the Inference Pipeline

_10.1 Distribution Type Selection Criteria
_10.2 Understanding Inference Distribution Types
_10.3 Comparison of Monolithic and Microservice Architectures
_10.4 Exploring LLM Twin's Inference Pipeline Deployment Strategies
_10.5 Deploying the LLM Twin Service
_10.6 Autoscaling to handle spikes in usage
_summation
_References

CHAPTER 11 MLOps and LLMOps

_11.1 DevOps, MLOps, LLMOps
_11.2 Deploying the LLM Twin Pipeline to the Cloud
_11.3 Applying LLMOps to LLM Twin
_summation
_References

APPENDIX MLOps Principles

_Principle 1: Automate or Operationalize
_Principle 2: Version Control
_Principle 3: Experimental Tracking
_Principle 4: Test
Principle 5: Monitoring
Principle 6: Reproducibility

Detailed image
Detailed Image 1

Publisher's Review
Design, implement, and learn your own AI
Everything about RAG, fine-tuning (LoRA·QLoRA), FastAPI, and LLMOps


ChatGPT is available to everyone, but it's not "tailor-made" for everyone.
Generic writing, long-winded answers, and inconsistent output are not what we want from AI.
This book goes beyond simple model calls and guides you through the entire process of developing a practical LLM system by implementing your own digital AI character, "LLM Twin."
From web scraping to RAG pipeline design, fine-tuning with LoRA and QLoRA, inference optimization, and cloud-based LLMOps, this book provides a hands-on project roadmap for developing end-to-end LLM applications.

In this process, readers will experience everything from data design, infrastructure configuration, and deployment strategies required to complete a real-world, production-grade system.
You can collect data from various sites like Medium, Substack, and GitHub, load it into MongoDB, optimize search performance using Qdrant, and even build microservices using RESTful APIs based on FastAPI.
This book goes beyond simply explaining complex LLM techniques, but instead unfolds them in a form that can be directly applied in practice. It is a practical guide suited to an era where we are moving beyond simply "utilizing" AI to "creating" AI ourselves.
It will serve as a definitive guide for developers, AI engineers, and technology leaders seeking to perfect their own LLM system.

Who is this book for?

● LLM Developer: Developers who want to go beyond using ChatGPT and create their own AI systems.
● AI Engineer: Professionals who want to practice the latest techniques such as RAG, LoRA, and QLoRA.
● ML System Engineer: Someone who wants to reliably deploy and operate AI services based on LLMOps
● Technology Leader: Team leaders who want to systematically learn the entire process of building an LLM, from data design to deployment.

What do you mainly cover?

● LLM application architecture design: Planning and system configuration of the personalized AI character 'LLM Twin'
● Data collection and preprocessing: Web scraping + MongoDB/Qdrant-based data storage and retrieval
● RAG Pipeline Development: Advanced Architecture Design and Document-Based Directive-Response Implementation
Supervised Learning Fine-Tuning (SFT): Generating a Directive Dataset + Utilizing LoRA and QLoRA
● Direct Preference Optimization (DPO): Fine-tuning sorting based on user preferences
● Model Evaluation and Tuning: LLM·RAG Performance Measurement and TwinLlama Experiments
● Inference optimization: Improving real-time inference performance through quantization, parallel processing, etc.
● LLM application deployment: FastAPI server implementation + auto-scaling-based deployment
● LLMOps Practical Application: Operational Automation Strategies, Including Version Management and Monitoring
GOODS SPECIFICS
- Date of issue: May 2, 2025
- Page count, weight, size: 508 pages | 918g | 183*235*22mm
- ISBN13: 9791169213806
- ISBN10: 1169213804

You may also like

카테고리