{"product_id":"110165","title":"Developing Practical AI Applications Using LLM ","description":"\u003ccenter\u003e\u003cdiv style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/tmgdisk01.cafe24.com\/images\/vs\/4172\/sv\/3jYEEOGwj3zG8SP2YwPWnS1C63zXnZ.png?v=1765095798\" style=\"max-width:100%;max-height:10px\"\u003e\u003c\/div\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\n\n\u003cdiv style=\"width:95%\"\u003e\n\n\u003cdiv style=\"text-align:center;font-size:30px;font-weight:bolder;line-height:1.6em\"\u003e Developing Practical AI Applications Using LLM \u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"border-bottom:1px;border-bottom-style:dotted;border-color:;padding-bottom:20px\"\u003e\u003ccenter\u003e\u003ctable align=\"center\" width=\"100%\"\u003e\u003ctbody style=\"border:0px\"\u003e\n\n\u003ctr\u003e\u003ctd align=\"center\" style=\"line-height:1.2em;text-align:center;font-size:18px;color:black;font-weight:bold;padding-bottom:20px;\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\u003ctr\u003e\u003ctd style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/image.yes24.com\/goods\/129081594\/XL\" style=\"max-width:100%;height:auto\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\n\u003c\/tbody\u003e\u003c\/table\u003e\u003c\/center\u003e\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;{split_style6}padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e Description \u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;word-break:break-all;font-size:14px;line-height:1.6em;\"\u003e\n\n\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eBook Introduction\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\u003cdiv\u003e \u003cb\u003eFrom Transformer architecture to RAG development, model training, deployment, optimization, and operation.\u003cbr\u003e Everything You Need to Know About AI Application Development with RamaIndex and LLM\u003c\/b\u003e\u003cbr\u003e \u003cbr\u003eThis book starts with the basic architecture of LLM, tames LLM according to application requirements, makes it lightweight to operate in limited computing environments, and lays the foundation for smooth service delivery. Then, it explains step-by-step how to create a representative LLM application called RAG.\u003cbr\u003e It doesn't end there. It covers methods for overcoming challenges encountered in actual operations, as well as advanced topics like multimodality and agents. It explains the essential development knowledge required for LLM students from both theoretical and practical perspectives, making it a valuable resource for developers seeking to adapt to the new paradigm.\u003cbr\u003e\n\u003c\/div\u003e\u003c\/div\u003e\n\u003cdiv\u003e\u003cul\u003e\u003cli\u003e You can preview some of the book's contents.\u003cbr\u003e \u003cspan\u003ePreview\u003c\/span\u003e\n\n\u003c\/li\u003e\u003c\/ul\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eindex\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e \u003cb\u003e[Part 1] Laying the Foundational Framework for the LLM\u003cbr\u003e\u003cbr\u003e Chapter 1 LLM Map\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 1.1 Deep Learning and Language Modeling \u003cbr\u003e__1.1.1 Deep learning that extracts data features on its own\u003cbr\u003e __1.1.2 Embedding: How Deep Learning Models Represent Data\u003cbr\u003e __1.1.3 Language Modeling: How Deep Learning Models Learn Language\u003cbr\u003e 1.2 How the Language Model Became ChatGPT\u003cbr\u003e __1.2.1 From RNN to Transformer Architecture\u003cbr\u003e __1.2.2 The relationship between model size and performance as seen in the GPT series\u003cbr\u003e __1.2.3 The emergence of ChatGPT\u003cbr\u003e 1.3 The era of LLM applications begins\u003cbr\u003e __1.3.1 LLM that Revolutionized the Way We Use Knowledge\u003cbr\u003e __1.3.2 sLLM: Building Smaller, More Efficient Models\u003cbr\u003e __1.3.3 Techniques for more efficient learning and inference\u003cbr\u003e __1.3.4 Augmented Search Generation (RAG) Technology to Address the Hallucination Phenomenon of LLM\u003cbr\u003e 1.4 The Future of LLM: Expanding Perception and Action\u003cbr\u003e 1.5 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 2: A Look at the Transformer Architecture, the Core of LLM\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 2.1 What is Transformer Architecture?\u003cbr\u003e 2.2 Converting text to embeddings\u003cbr\u003e __2.2.1 Tokenization\u003cbr\u003e __2.2.2 Converting to token embeddings\u003cbr\u003e __2.2.3 Positional Encoding\u003cbr\u003e 2.3 Understanding Attention \u003cbr\u003e__2.3.1 How People Read and Attention\u003cbr\u003e __2.3.2 Understanding queries, keys, and values\u003cbr\u003e __2.3.3 Attention in Code\u003cbr\u003e __2.3.4 Multi-head attention\u003cbr\u003e 2.4 Normalization and feedforward layers\u003cbr\u003e __2.4.1 Understanding Layer Normalization\u003cbr\u003e __2.4.2 Feedforward layer\u003cbr\u003e 2.5 Encoder\u003cbr\u003e 2.6 Decoder\u003cbr\u003e 2.7 Architectures utilizing transformers such as BERT, GPT, and T5\u003cbr\u003e BERT using __2.7.1 encoder\u003cbr\u003e __2.7.2 GPT using decoder\u003cbr\u003e __2.7.3 BART, T5 using both encoder and decoder\u003cbr\u003e 2.8 Main pre-learning mechanisms\u003cbr\u003e __2.8.1 Causal Language Modeling\u003cbr\u003e __2.8.2 Mask Language Modeling\u003cbr\u003e 2.9 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 3: Hugging Face Transformer Library for Handling Transformer Models\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 3.1 What is a Hugging Face Transformer?\u003cbr\u003e 3.2 Exploring the Hugging Face Hub\u003cbr\u003e __3.2.1 Model Hub\u003cbr\u003e __3.2.2 Dataset Hub\u003cbr\u003e __3.2.3 A space where model demos can be made public and available for use\u003cbr\u003e 3.3 Learning how to use the Hugging Face library\u003cbr\u003e __3.3.1 Using the Model\u003cbr\u003e __3.3.2 Using the Tokenizer \u003cbr\u003e__3.3.3 Using the dataset\u003cbr\u003e 3.4 Training the Model\u003cbr\u003e __3.4.1 Data Preparation\u003cbr\u003e __3.4.2 Training using the Trainer API\u003cbr\u003e __3.4.3 Training without using the trainer API\u003cbr\u003e __3.4.4 Uploading the trained model\u003cbr\u003e 3.5 Model Inference\u003cbr\u003e __3.5.1 Inference using pipelines\u003cbr\u003e __3.5.2 Direct Inference\u003cbr\u003e 3.6 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 4: Building a Model That Obeys\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 4.1 Passing the Coding Test: Pretraining and Fine-tuning the Map\u003cbr\u003e __4.1.1 Learning Coding Concepts: Prerequisites for the LLM\u003cbr\u003e __4.1.2 Practice Problem: Fine-tuning the Map\u003cbr\u003e __4.1.3 Conditions for a good instruction dataset\u003cbr\u003e 4.2 Improving Code Readability with a Scoring Model\u003cbr\u003e __4.2.1 Creating a Scoring Model Using a Preferred Dataset\u003cbr\u003e __4.2.2 Reinforcement Learning: Towards High Code Readability Scores\u003cbr\u003e __4.2.3 PPO: Avoiding Compensation Hacks\u003cbr\u003e __4.2.4 RLHF: Cool, but if you can avoid it…\u003cbr\u003e 4.3 Is reinforcement learning really necessary?\u003cbr\u003e __4.3.1 Rejection Sampling: What if we simply use the data with the highest scores? \u003cbr\u003e__4.3.2 DPO: Training on Preferred Datasets Directly\u003cbr\u003e __4.3.3 Models trained using DPO\u003cbr\u003e 4.4 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003e[Part 2: Taming the LLM]\u003cbr\u003e\u003cbr\u003e Chapter 5: GPU-Efficient Learning\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 5.1 Examining Data Uploaded to the GPU\u003cbr\u003e __5.1.1 Data types of deep learning models\u003cbr\u003e __5.1.2 Reducing Model Size with Quantization\u003cbr\u003e __5.1.3 Disassembling GPU Memory\u003cbr\u003e 5.2 Efficiently Utilizing a Single GPU\u003cbr\u003e __5.2.1 Gradient Accumulation\u003cbr\u003e __5.2.2 Gradient Checkpointing\u003cbr\u003e 5.3 Distributed Learning and ZeRO\u003cbr\u003e __5.3.1 Distributed Learning\u003cbr\u003e __5.3.2 Reducing Redundant Storage in Data Parallelism (ZeRO)\u003cbr\u003e 5.4 Efficient Learning Method (PEFT): LoRA\u003cbr\u003e __5.4.1 LoRA learning by reconfiguring only a portion of the model parameters\u003cbr\u003e __5.4.2 Viewing LoRA Settings\u003cbr\u003e __5.4.3 Using LoRA Learning with Code\u003cbr\u003e 5.5 Efficient Learning Method (PEFT): QLoRA\u003cbr\u003e __5.5.1 4-bit quantization and second-order quantization\u003cbr\u003e __5.5.2 Page Optimizer\u003cbr\u003e __5.5.3 Using QLoRA Models with Code\u003cbr\u003e 5.6 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 6: Studying sLLM\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 6.1 Text2SQL Dataset\u003cbr\u003e __6.1.1 Representative Text2SQL datasets \u003cbr\u003e__6.1.2 Korean dataset\u003cbr\u003e __6.1.3 Using synthetic data\u003cbr\u003e 6.2 Preparing the Performance Evaluation Pipeline\u003cbr\u003e __6.2.1 Text2SQL Evaluation Method\u003cbr\u003e __6.2.2 Building an Evaluation Dataset\u003cbr\u003e __6.2.3 SQL Generation Prompt\u003cbr\u003e __6.2.4 GPT-4 Evaluation Prompt and Code Preparation\u003cbr\u003e 6.3 Hands-on: Performing Fine-Tuning\u003cbr\u003e __6.3.1 Evaluating the Base Model\u003cbr\u003e __6.3.2 Performing fine tuning\u003cbr\u003e __6.3.3 Training Data Cleaning and Fine-Tuning\u003cbr\u003e __6.3.4 Change the basic model\u003cbr\u003e __6.3.5 Model Performance Comparison\u003cbr\u003e 6.4 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 7: Making the Model Lighter\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 7.1 Understanding Language Model Inference\u003cbr\u003e __7.1.1 How Language Models Generate Language\u003cbr\u003e __7.1.2 KV cache to reduce redundant operations\u003cbr\u003e __7.1.3 GPU Architecture and Optimal Batch Size\u003cbr\u003e __7.1.4 Reducing KV cache memory\u003cbr\u003e 7.2 Reducing model size with quantization\u003cbr\u003e __7.2.1 Bits and Bites\u003cbr\u003e __7.2.2 GPTQ\u003cbr\u003e __7.2.3 AWQ\u003cbr\u003e 7.3 Using Knowledge Distillation\u003cbr\u003e 7.4 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 8 Serving sLLM\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 8.1 Efficient Deployment Strategies\u003cbr\u003e __8.1.1 General layout (static layout)\u003cbr\u003e __8.1.2 Dynamic Deployment\u003cbr\u003e __8.1.3 Continuous Batch \u003cbr\u003e8.2 Efficient transformer operations\u003cbr\u003e __8.2.1 Flash Attention\u003cbr\u003e __8.2.2 Flash Attention 2\u003cbr\u003e __8.2.3 Relative Position Encoding\u003cbr\u003e 8.3 Efficient Inference Strategies\u003cbr\u003e __8.3.1 Kernel Fusion\u003cbr\u003e __8.3.2 Page Attention\u003cbr\u003e __8.3.3 Speculative Decoding\u003cbr\u003e 8.4 Hands-on: LLM Serving Framework\u003cbr\u003e __8.4.1 Offline Serving\u003cbr\u003e __8.4.2 Online Serving\u003cbr\u003e 8.5 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003e[Part 3] Practical Application Development Using LLM\u003cbr\u003e\u003cbr\u003e Chapter 9: Developing LLM Applications\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 9.1 Augmented Search Generation (RAG)\u003cbr\u003e __9.1.1 Data Storage\u003cbr\u003e __9.1.2 Integrating search results into prompts\u003cbr\u003e __9.1.3 Practice: Implementing RAG with RamaIndex\u003cbr\u003e 9.2 LLM Cache\u003cbr\u003e __9.2.1 How LLM Cache Works\u003cbr\u003e __9.2.2 Hands-on: Implementing the OpenAI API Cache\u003cbr\u003e 9.3 Data Verification\u003cbr\u003e __9.3.1 Data Verification Method\u003cbr\u003e __9.3.2 Data Validation Practice\u003cbr\u003e 9.4 Data Logging\u003cbr\u003e __9.4.1 OpenAI API Logging\u003cbr\u003e __9.4.2 RamaIndex Logging\u003cbr\u003e 9.5 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 10: Compressing Data Meaning with Embedding Models\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 10.1 Understanding Text Embeddings\u003cbr\u003e __10.1.1 Advantages of sentence embedding\u003cbr\u003e __10.1.2 One-Hot Encoding\u003cbr\u003e __10.1.3 Back of Wars \u003cbr\u003e__10.1.4 TF-IDF\u003cbr\u003e __10.1.5 Word to Back\u003cbr\u003e 10.2 Sentence embedding method\u003cbr\u003e __10.2.1 Two ways to compute relationships between sentences\u003cbr\u003e __10.2.2 Bi-encoder model structure\u003cbr\u003e __10.2.3 Generating Text and Image Embeddings with Sentence-Transformers\u003cbr\u003e __10.2.4 Comparing Open Source and Commercial Embedding Models\u003cbr\u003e 10.3 Hands-on: Implementing Semantic Search\u003cbr\u003e __10.3.1 Implementing semantic search\u003cbr\u003e __10.3.2 Using the Sentence-Transformers Model in RamaIndex\u003cbr\u003e 10.4 Combining Search Methods to Improve Performance\u003cbr\u003e __10.4.1 Keyword Search Method: BM25\u003cbr\u003e __10.4.2 Understanding Mutual Rank Combinations\u003cbr\u003e 10.5 Hands-on: Implementing Hybrid Search\u003cbr\u003e __10.5.1 Implementing BM25\u003cbr\u003e __10.5.2 Implementing Mutual Rank Combinations\u003cbr\u003e __10.5.3 Implementing Hybrid Search\u003cbr\u003e 10.6 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 11: Building Embedding Models Tailored to Your Data: Improving RAG\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 11.1 Two Ways to Improve Search Performance\u003cbr\u003e 11.2 Creating a Language Model as an Embedding Model\u003cbr\u003e __11.2.1 Contrastive Learning\u003cbr\u003e __11.2.2 Practice: Preparing for Learning \u003cbr\u003e__11.2.3 Practice: Training an Embedding Model with Similar Sentence Data\u003cbr\u003e 11.3 Fine-tuning the embedding model\u003cbr\u003e __11.3.1 Practice: Preparing for Learning\u003cbr\u003e __11.3.2 Fine-tuning using MNR loss\u003cbr\u003e 11.4 Reordering rankings to improve search quality\u003cbr\u003e 11.5 Implementing an Improved RAG with Bi-Encoder and Cross-Encoder\u003cbr\u003e __11.5.1 Searching with the default embedding model\u003cbr\u003e __11.5.2 Searching with a Fine-Tuned Embedding Model\u003cbr\u003e __11.5.3 Combining a Fine-Tuned Embedding Model with a Cross-Encoder\u003cbr\u003e 11.6 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 12: Extending to Vector Databases: Implementing RAG\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 12.1 What is a vector database?\u003cbr\u003e __12.1.1 Deep Learning and Vector Databases\u003cbr\u003e __12.1.2 Understanding the Vector Database Terrain\u003cbr\u003e 12.2 How Vector Databases Work\u003cbr\u003e __12.2.1 KNN Search and Its Limitations\u003cbr\u003e __12.2.2 What is ANN search?\u003cbr\u003e __12.2.3 Explorable Small World (NSW)\u003cbr\u003e __12.2.4 Hierarchical Structure\u003cbr\u003e 12.3 Exercise: Understanding the Key Parameters of the HNSW Index\u003cbr\u003e __12.3.1 Understanding the parameter m \u003cbr\u003e__12.3.2 Understanding the ef_construction parameter\u003cbr\u003e __12.3.3 Understanding the ef_search parameter\u003cbr\u003e 12.4 Hands-on: Implementing Vector Search with Pinecone\u003cbr\u003e __12.4.1 How to use the Pinecone client\u003cbr\u003e __12.4.2 Changing the vector database in the RamaIndex\u003cbr\u003e 12.5 Hands-on: Implementing Multimodal Search with Pinecone\u003cbr\u003e __12.5.1 Dataset\u003cbr\u003e __12.5.2 Practice Flow\u003cbr\u003e __12.5.3 Generating Image Descriptions with GPT-4o\u003cbr\u003e __12.5.4 Save Prompt\u003cbr\u003e __12.5.5 Image Embedding Search\u003cbr\u003e __12.5.6 Creating an image with DALL-E 3\u003cbr\u003e 12.6 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 13: Running an LLM\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 13.1 MLOps\u003cbr\u003e __13.1.1 Data Management\u003cbr\u003e __13.1.2 Experimental Management\u003cbr\u003e __13.1.3 Model Repository\u003cbr\u003e __13.1.4 Model Monitoring\u003cbr\u003e 13.2 What is different about LLMOps?\u003cbr\u003e __13.2.1 Choosing between a commercial and open-source model\u003cbr\u003e __13.2.2 Changes in model optimization methods\u003cbr\u003e __13.2.3 Difficulties in LLM Evaluation\u003cbr\u003e 13.3 Evaluating the LLM\u003cbr\u003e __13.3.1 Quantitative indicators\u003cbr\u003e __13.3.2 Evaluation using benchmark datasets\u003cbr\u003e __13.3.3 How people directly evaluate\u003cbr\u003e __13.3.4 Evaluation through LLM \u003cbr\u003e__13.3.4 RAG Evaluation\u003cbr\u003e 13.4 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003e[Part 4] Multimodality, Agents, and the Future of LLM\u003cbr\u003e\u003cbr\u003e Chapter 14 Multimodal\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e LLM 14.1 What is a Multimodal LLM?\u003cbr\u003e __14.1.1 Components of a Multimodal LLM\u003cbr\u003e __14.1.2 Multimodal LLM Learning Course\u003cbr\u003e 14.2 Model for linking images and text: CLIP\u003cbr\u003e __14.2.1 What is the CLIP model?\u003cbr\u003e __14.2.2 CLIP model learning method\u003cbr\u003e __14.2.3 Utilization of CLIP model and excellent performance\u003cbr\u003e __14.2.4 Using the CLIP Model Directly\u003cbr\u003e 14.3 A model for generating images from text: DALL-E\u003cbr\u003e __14.3.1 Diffusion Model Principle\u003cbr\u003e __14.3.2 DALL-E model\u003cbr\u003e 14.4 LLaVA\u003cbr\u003e __14.4.1 LLaVA's training data\u003cbr\u003e __14.4.2 LLaVA model structure\u003cbr\u003e __14.4.3 LLaVA 1.5\u003cbr\u003e __14.4.4 LLaVA NeXT\u003cbr\u003e 14.5 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 15 LLM Agent\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 15.1 What is an Agent?\u003cbr\u003e __15.1.1 Agent Components\u003cbr\u003e __15.1.2 The Agent's Brain\u003cbr\u003e __15.1.3 Agent's Senses\u003cbr\u003e __15.1.4 Agent Behavior\u003cbr\u003e 15.2 Types of Agent Systems\u003cbr\u003e __15.2.1 Single Agent\u003cbr\u003e __15.2.2 User-Agent Interaction\u003cbr\u003e __15.2.3 Multi-Agent\u003cbr\u003e 15.3 Evaluating Agents \u003cbr\u003e15.4 Hands-on: Implementing an Agent\u003cbr\u003e __15.4.1 Basic AutoGen Usage\u003cbr\u003e __15.4.2 RAG Agent\u003cbr\u003e __15.4.3 Multimodal Agent\u003cbr\u003e 15.5 Summary\u003cbr\u003e\u003cbr\u003e \u003cb\u003eChapter 16: New Architecture\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 16.1 Advantages and Disadvantages of Existing Architectures\u003cbr\u003e 16.2 SSM\u003cbr\u003e __16.2.1 S4\u003cbr\u003e 16.3 Selection Mechanism\u003cbr\u003e 16.4 Mamba\u003cbr\u003e __16.4.1 Mamba Performance\u003cbr\u003e __16.4.2 Comparison with existing architecture\u003cbr\u003e Mamba in 16.5 code\u003cbr\u003e\u003cbr\u003e \u003cb\u003eAppendix | Preparation for the Practice\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e A.1 How to Use Google Colab\u003cbr\u003e A.2 Hugging Face Token\u003cbr\u003e A.3 OpenAI Token\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eDetailed image\u003c\/b\u003e \u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\u003cdiv\u003e\u003cimg src=\"https:\/\/image.yes24.com\/momo\/TopCate4618\/MidCate010\/461793209.jpg\" border=\"0\" alt=\"Detailed Image 1\"\u003e\u003c\/div\u003e\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003ePublisher's Review\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e \u003cb\u003e| What this book covers |\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e - Transformer architecture, the core of LLM\u003cbr\u003e - How to create a ChatGPT: Map fine-tuning and RLHF\u003cbr\u003e - Learn more about open source LLM with your own data\u003cbr\u003e - Lightweight model for LLM application operation\u003cbr\u003e - Implementation and improvement of RAG using Rama Index\u003cbr\u003e - Multimodal LLM that processes both images and voices\u003cbr\u003e - Agent architecture combining long-term memory and tools in LLM\u003cbr\u003e\u003cbr\u003e \u003cb\u003e| Target audience for this book |\u003c\/b\u003e\u003cbr\u003e \u003cbr\u003e- Developers who want to develop AI applications using LLM\u003cbr\u003e - Developers who are curious about the principles and underlying technologies of the model rather than simply utilizing the LLM API.\u003cbr\u003e - Students and job seekers who want to become AI engineers\u003cbr\u003e - Graduate students who want to organize LLM-related papers and technologies in a short period of time\u003cbr\u003e\u003cbr\u003e \u003cb\u003e| Download the GitHub practice code |\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e The practice code can be found in the book's GitHub repository (https:\/\/github.com\/onlybooks\/llm).\u003cbr\u003e Code from GitHub can be leveraged in Google Colab in two ways:\u003cbr\u003e\u003cbr\u003e 1.\u003cbr\u003e Upload from local: Clone the code from GitHub to your local environment or download it as a compressed file, then open the notebook file (ipynb) in the practice folder you want to proceed with in Google Colab.\u003cbr\u003e You can proceed with the wet.\u003cbr\u003e 2. \u003cbr\u003eOpen with GitHub URL: When you select Open Notebook in Google Colab (Ctrl+O), you can open the practice notebook via the code URL in the GitHub tab among various ways to open the notebook.\u003cbr\u003e\u003cbr\u003e \u003cb\u003e| Code execution environment for this book |\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e The exercises in this book are run on Google Colab.\u003cbr\u003e Google Colab is a notebook execution environment provided by Google that can be run in a browser with a UI similar to Python's Jupyter Notebook.\u003cbr\u003e Google Colab also provides T4 GPU (16GB) for free use.\u003cbr\u003e The free version of Google Colab has a 12-hour runtime limit and may disconnect if left unused for extended periods.\u003cbr\u003e\u003cbr\u003e \u003cb\u003e[Author's Note]\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e As of July 2024, when this article was written, the keywords in the AI ​​and LLM market can be said to be 'multi-modal, agent, and on-device AI.' \u003cbr\u003eMultimodal refers to an AI model that processes various types of data, such as text, images, voice, and video, while an agent refers to a more advanced system in which an AI model has long-term memory and utilizes various tools, such as internet searches and code execution, to solve users' problems.\u003cbr\u003e On-device AI refers to AI models running directly on the user's device, rather than in the cloud or on high-performance servers.\u003cbr\u003e The advantage is that users can utilize AI models without worrying about their personal information being leaked, as their information does not leave the device.\u003cbr\u003e This major trend in the LLM market can be confirmed through models and projects recently released by leading AI companies.\u003cbr\u003e\u003cbr\u003e In May 2024, OpenAI unveiled GPT-4o, a new language model that can see, hear, and speak. \u003cbr\u003eA day later, Google DeepMind unveiled Project Astra, a multimodal agent that does almost the same thing.\u003cbr\u003e In June 2024, Anthropic released its Claude 3.5 Sonnet model, which outperformed OpenAI's GPT-4o.\u003cbr\u003e Sonnet is the mid-level model name of Anthropic.\u003cbr\u003e This makes us anticipate how excellent the upcoming high-performance model, the Claude 3.5 Opus, will be.\u003cbr\u003e\u003cbr\u003e In June 2024, Apple announced a new name for the AI ​​capabilities running on its devices: Apple Intelligence.\u003cbr\u003e Since the release of ChatGPT, Apple has shown no significant movement in AI research and development, leading many to believe that Apple is losing its leadership in the AI ​​era. \u003cbr\u003eHowever, Apple has shown confidence by using a provocative name that means AI, which can also be read as Apple Intelligence, and the market is also showing great expectations.\u003cbr\u003e\u003cbr\u003e To understand the key keywords mentioned so far, you need to understand the latest AI models themselves, including large language models, and how to utilize them.\u003cbr\u003e\u003cbr\u003e This book covers both the models themselves and their application, helping readers keep up with recent trends in the AI ​​market.\u003cbr\u003e The first two parts of the book (Chapters 1-8) introduce the principles of LLM and the model itself, including how to learn the model and how to make inferences, to help you understand it in depth.\u003cbr\u003e Part 3 of the book (Chapters 9-13) explores the components needed to develop applications using LLM and Retrieval Augmented Generation (RAG). \u003cbr\u003eFinally, Part 4 (Chapters 14-16) introduces multimodality, agents, and newly researched LLM architectures, providing a glimpse into the near future of LLM.\u003cbr\u003e\u003cbr\u003e \u003cb\u003e[Editor's Note]\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Before discussing LLM application development, it is necessary to consider why LLM applications are so appealing to people.\u003cbr\u003e Numerous technologies have emerged and disappeared over the years, but the major technologies still in use today have the distinction of having significantly advanced development or usage paradigms.\u003cbr\u003e So, to determine whether LLM applications will become mainstream and become a vital toolbox in developers' brains, it's best to first look at what changes are taking place.\u003cbr\u003e\u003cbr\u003e\u003cbr\u003e Let's start with changes on the user side.\u003cbr\u003e Even now, people need to keep search in mind to acquire knowledge. \u003cbr\u003eWhen a user enters a keyword or combination of words into a search engine, a list of matching documents magically appears.\u003cbr\u003e In the beginning, people meticulously compiled the list one by one like a phone book to improve quality, but soon the era of information overload arrived, and automation inevitably took place. The most representative search engine, the AltaVista series, did not have an excellent algorithm for ranking search results, so people had to manually look through numerous lists.\u003cbr\u003e Then, Google, which we still rely on to this day, came along and greatly improved search quality so that you could find what you wanted just by looking at the first page.\u003cbr\u003e\u003cbr\u003e\u003cbr\u003e But recently, the focus has shifted from searching to asking questions. \u003cbr\u003eIn today's fast-paced, busy society, even the act of opening a list one by one has become burdensome, and the quality of search results is continuously declining compared to the past due to cases of misuse under the guise of search engine optimization. Therefore, it can be seen as matching the desire to ask a question and then get the answer I want.\u003cbr\u003e In the case of ChatGPT or Google Gemini, multi-modality is supported, so you can ask and receive questions by combining images and text.\u003cbr\u003e\u003cbr\u003e Next, we need to look at the changes in machine learning and deep learning.\u003cbr\u003e Before the advent of generative AI and large language models (LLMs), most machine learning and deep learning models were domain-specific.\u003cbr\u003e In other words, research and development has been conducted to properly solve narrow-scope problems so that predictions and insights can be made by utilizing data accumulated in the relevant domain. \u003cbr\u003eOf course, when the world of deep learning began, big data, cloud, and GPU infrastructure already existed, but there was no technological or economic breakthrough to expand them for general purposes.\u003cbr\u003e For example, in the field of images, various deep learning models have emerged, starting with digit judgment using the MNIST dataset and various class judgments using the CIFAR-10 dataset, and even YOLO, which can identify the location of objects and judge classes in real time, but they are all limited to image classification.\u003cbr\u003e\u003cbr\u003e\u003cbr\u003e However, with the emergence of generative AI, it is evolving into a general-purpose model capable of performing a variety of tasks.\u003cbr\u003e For example, ChatGPT is quite good at answering not only general questions but also specialized questions such as medical or legal ones, and advanced models such as GPT-4V or 4o that support multi-modality can also answer questions about tables or images. \u003cbr\u003eIt is also considered a versatile tool because it can generate Python code at runtime and display the results of execution in an isolated environment when complex calculations or graph output are required.\u003cbr\u003e Generative AI is demonstrating diverse use cases, including Q\u0026amp;A, translation, classification, summarization, analysis, stylistic modification, and sentiment analysis, and is evolving beyond information and decision support to the level of autonomous reasoning.\u003cbr\u003e\u003cbr\u003e\u003cbr\u003e Finally, we can't leave out the changes in development itself, which developers are likely to be most interested in.\u003cbr\u003e Early programming revolved around logic. Functional programming languages ​​like LISP challenged human computational abilities based on mathematical theory. Procedural methods were introduced through Fortran, Pascal, and the C programming language. Object-oriented techniques were introduced through C++ and Java. \u003cbr\u003eMeanwhile, advancements have been centered around data. SQL made it easy to manipulate structured data in a tabular format in a structured way, and with the advent of the big data era, it became possible to handle not only structured data but also semi-structured and unstructured data.\u003cbr\u003e\u003cbr\u003e\u003cbr\u003e Meanwhile, with the recent emergence of generative AI, another wave of change is being detected.\u003cbr\u003e A new method has emerged: using a language more closely related to people, rather than a language more closely related to computers, to implement business logic! This new technique, also known as prompt engineering, orchestrates LLM inputs to achieve desired results and adjusts the output to achieve the desired outcome. \u003cbr\u003eSince prompts are closer to the language used by humans, it is difficult to guarantee accuracy, but in return, flexibility and extensibility are gained, which has created an opportunity to easily implement business logic that was previously quite difficult.\u003cbr\u003e\u003cbr\u003e As generative AI emerges, it brings with it a number of changes, and developers must prepare to adapt.\u003cbr\u003e If you are developing an application with the existing web browser (or app) - WAS (web application server) - database (relational or NoSQL) structure, you need to identify the use cases of LLM according to corporate requirements and change the architecture accordingly. \u003cbr\u003eA technology of particular note here is Augmented Search Generation (RAG), which has recently been attracting significant attention in the corporate world. RAG is an application that utilizes a technology called embedding various documents and data within a company to build a knowledge base as a vector database. It then extracts the document fragments most relevant to the user's question from the knowledge base and has the LLM summarize them.\u003cbr\u003e\u003cbr\u003e \u003cbr\u003eWhile traditional search methods based on algorithms such as TF\/IDF or BM25 encode words (keywords) in a sentence using sparse vectors, the semantic search used in RAG uses dense vectors to encode the abstract meaning and relationships of words, so it can be seen as a great match with LLM, which has strengths in language processing. RAG can be seen as a good example of the concept of information entropy advocated by Claude Shannon in the 1950s being realized in earnest, and it appropriately combines the two major technologies of embedding and LLM to achieve the best performance in order to process information containing various meanings, so developers will be able to gain a lot of inspiration just by looking at this technology ecosystem itself.\u003cbr\u003e\u003cbr\u003e \u003cbr\u003eThis book starts with the basic architecture of LLM, tames LLM to meet application requirements, makes it lightweight for operation in limited computing environments, and lays the foundation for smooth service delivery. It then explains step-by-step how to create a representative LLM application called RAG.\u003cbr\u003e It doesn't end here, but covers how to overcome difficulties encountered in actual operation, as well as advanced topics such as multi-modality and agents.\u003cbr\u003e In other words, it explains the development knowledge essential for the LLM era from both theoretical and practical perspectives, so it will be a welcome relief to developers seeking to adapt to the new paradigm.\u003cbr\u003e I highly recommend this to all developers who are constantly working on research and development.\u003cbr\u003e\u003cbr\u003e - Jaeho Park \/ Operator of the blog \"Computer vs. Book\" and translator of \"Clean Code: Now Python\" (Bookman, 2022) \u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e GOODS SPECIFICS\u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;font-size:14px;line-height:1.6em;\"\u003e\n\n \u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e- \u003cstrong\u003eDate of issue:\u003c\/strong\u003e July 25, 2024\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003ePage count, weight, size:\u003c\/strong\u003e 556 pages | 1,030g | 184*240*28mm\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN13:\u003c\/strong\u003e 9791189909703\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN10:\u003c\/strong\u003e 1189909707 \u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cspan\u003e\u003c\/span\u003e\n\n\u003c\/center\u003e\n\n\n\u003c\/center\u003e","brand":"LIBRAIRIE COREENNE","offers":[{"title":"Default Title","offer_id":43893684568106,"sku":"110165","price":46.0,"currency_code":"EUR","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0683\/2750\/5962\/files\/5af9db55b3df68eb95500569ca558635.jpg?v=1765411655","url":"https:\/\/librairie.coreenne.fr\/en\/products\/110165","provider":"LIBRAIRIE COREENNE","version":"1.0","type":"link"}