
LLM Service Design and Optimization
Description
Book Introduction
LLM Optimization Strategies to Enhance the Competitiveness of Generative AI
Advances in AI and machine learning have led to a surge in interest in large language models (LLMs), but their high cost has hindered many companies from adopting them.
This book presents an efficient approach to building and deploying LLMs at low cost.
Learn how to effectively minimize costs while minimizing performance degradation at each stage: model selection, prompt engineering, fine-tuning, and deployment.
Provides practical and technical knowledge required to implement generative AI applications such as search systems and AI agents.
Let's strengthen the competitiveness of generative AI services by exploring inference optimization techniques such as model quantization and scaling, as well as ways to reduce infrastructure costs.
Advances in AI and machine learning have led to a surge in interest in large language models (LLMs), but their high cost has hindered many companies from adopting them.
This book presents an efficient approach to building and deploying LLMs at low cost.
Learn how to effectively minimize costs while minimizing performance degradation at each stage: model selection, prompt engineering, fine-tuning, and deployment.
Provides practical and technical knowledge required to implement generative AI applications such as search systems and AI agents.
Let's strengthen the competitiveness of generative AI services by exploring inference optimization techniques such as model quantization and scaling, as well as ways to reduce infrastructure costs.
- You can preview some of the book's contents.
Preview
index
CHAPTER 1 LLM Foundations
_1.1 Generative AI Applications and LLM
_1.2 The Path to Commercialization of Generative AI Applications
_1.3 The Importance of Cost Optimization
_1.4 Summary
CHAPTER 2: Tuning Techniques for Cost Optimization
_2.1 Fine Tuning and Customizing
_2.2 Parameter Efficient Fine Tuning (PEFT)
_2.3 Impact of PEFT on Cost and Performance
_2.4 Summary
CHAPTER 3: Inference Techniques for Cost Optimization
_3.1 Introduction to Inference Techniques
_3.2 Prompt Engineering
_3.3 Caching using vector stores
_3.4 Chain for managing long documents
_3.5 Text Summary
_3.6 Batching Prompts for Efficient Inference
_3.7 Model Optimization Methods
_3.8 Parameter Efficient Fine Tuning (PEFT)
_3.9 Cost and Performance Impacts
_3.10 Summary
CHAPTER 4 MODEL CHOICE AND ALTERNATIVES
_4.1 The Importance of Model Selection
_4.2 Efficient compact model
_4.3 Successful Small Model Cases
_4.4 Domain-Specific Model
_4.5 Performance of prompts using a general model
_4.6 Summary
CHAPTER 5 Infrastructure and Deployment Tuning Strategies
_5.1 Tuning Strategy
_5.2 Hardware Utilization and Deployment Tuning
_5.3 Inference Acceleration Tool
_5.4 Monitoring and Observability
_5.5 Summary
CHAPTER 6: The Key to Successful Generative AI Implementation
_6.1 Balance between performance and cost
_6.2 Future Trends in Generative AI Applications
_6.3 Summary
_1.1 Generative AI Applications and LLM
_1.2 The Path to Commercialization of Generative AI Applications
_1.3 The Importance of Cost Optimization
_1.4 Summary
CHAPTER 2: Tuning Techniques for Cost Optimization
_2.1 Fine Tuning and Customizing
_2.2 Parameter Efficient Fine Tuning (PEFT)
_2.3 Impact of PEFT on Cost and Performance
_2.4 Summary
CHAPTER 3: Inference Techniques for Cost Optimization
_3.1 Introduction to Inference Techniques
_3.2 Prompt Engineering
_3.3 Caching using vector stores
_3.4 Chain for managing long documents
_3.5 Text Summary
_3.6 Batching Prompts for Efficient Inference
_3.7 Model Optimization Methods
_3.8 Parameter Efficient Fine Tuning (PEFT)
_3.9 Cost and Performance Impacts
_3.10 Summary
CHAPTER 4 MODEL CHOICE AND ALTERNATIVES
_4.1 The Importance of Model Selection
_4.2 Efficient compact model
_4.3 Successful Small Model Cases
_4.4 Domain-Specific Model
_4.5 Performance of prompts using a general model
_4.6 Summary
CHAPTER 5 Infrastructure and Deployment Tuning Strategies
_5.1 Tuning Strategy
_5.2 Hardware Utilization and Deployment Tuning
_5.3 Inference Acceleration Tool
_5.4 Monitoring and Observability
_5.5 Summary
CHAPTER 6: The Key to Successful Generative AI Implementation
_6.1 Balance between performance and cost
_6.2 Future Trends in Generative AI Applications
_6.3 Summary
Detailed image

Publisher's Review
Now, the core of AI services is optimization!
Learn everything about LLM Service Design!
With the emergence of LLMs that deliver high performance with minimal investment, such as DeepSearch, a new keyword has emerged in the AI development process: optimization.
This book broadly covers the practical methodologies and theories needed to build high-performance AI services with efficient investments, from leveraging small models (SLMs), effective prompt engineering, fine-tuning, and quantization techniques.
Drawing on a variety of case studies and theories, we provide in-depth insights to startups, companies, and developers struggling with the costs of adopting AI technology.
We hope this will be of practical help to anyone who needs an LLM optimization strategy that reduces costs and improves performance.
Key Contents
● An effective technique to solve the problem of high computational cost of LLM
Fine-tuning, inference, and quantization techniques to create cost-effective generative AI services
● Alternative models such as small models and domain-specific models
Target audience
● Practical engineers who want to build, tune, and deploy efficient AI models
● Planners and decision makers who want to make a business evaluation of AI services
● Developers who want to learn about the overall technology of artificial intelligence models, including LLM
● Students and professors studying generative AI and LLM
Learn everything about LLM Service Design!
With the emergence of LLMs that deliver high performance with minimal investment, such as DeepSearch, a new keyword has emerged in the AI development process: optimization.
This book broadly covers the practical methodologies and theories needed to build high-performance AI services with efficient investments, from leveraging small models (SLMs), effective prompt engineering, fine-tuning, and quantization techniques.
Drawing on a variety of case studies and theories, we provide in-depth insights to startups, companies, and developers struggling with the costs of adopting AI technology.
We hope this will be of practical help to anyone who needs an LLM optimization strategy that reduces costs and improves performance.
Key Contents
● An effective technique to solve the problem of high computational cost of LLM
Fine-tuning, inference, and quantization techniques to create cost-effective generative AI services
● Alternative models such as small models and domain-specific models
Target audience
● Practical engineers who want to build, tune, and deploy efficient AI models
● Planners and decision makers who want to make a business evaluation of AI services
● Developers who want to learn about the overall technology of artificial intelligence models, including LLM
● Students and professors studying generative AI and LLM
GOODS SPECIFICS
- Date of issue: April 10, 2025
- Page count, weight, size: 296 pages | 183*235*12mm
- ISBN13: 9791169213646
- ISBN10: 1169213642
You may also like
카테고리
korean
korean