Skip to product information
Hands-on generative AI
Hands-on generative AI
Description
Book Introduction
Generative AI: Beyond Theory, into Practice
Master GenAI Practical Implementation in One Book with Hugging Face Core Developers


The author who developed the Hugging Face Core personally implements the latest technology to resolve the gap between theory and practice and the thirst for technology.
Beyond propulsion engineering, it explores the internal structure of transformers and diffusion models, allowing you to understand the core principles of generative AI and learn by implementing cutting-edge technologies yourself.
It covers the main structure and operation of generative AI, focusing on transformers and diffusion models, and provides an in-depth explanation of how multimodal models that generate images, text, and audio work and how to utilize them.

It organizes core concepts such as autoencoders, CLIP, and U-Net, and is structured so that you can learn both concepts and implementation through various project-centered exercises ranging from text generation, conditional image generation, and audio generation.
In particular, you can practice directly without complicated settings by utilizing the environment based on Hugging Face and Google Colab, and you can also implement the latest technologies such as Stable Diffusion, Dreambooth, and LoRA step by step.
It also introduces transfer learning techniques required for practical applications, from text classification, generation, and directive-based fine-tuning to Augmented Augmentation Generation (RAG) implementation, along with actual code. In addition to creative application examples such as inpainting, image editing, and control nets, it also examines the development trends of the latest generative AI technologies such as multimodality, 3D vision, and video generation.
For developers seeking to apply generative AI in practice, this book will serve as a comprehensive, systematic guide, covering everything from technical principles to implementation and application.

  • You can preview some of the book's contents.
    Preview
","
index
[Part 1: Utilizing the Open Model]

Chapter 1: Introduction to Generative Media
_1.1 Image Creation
_1.2 Text Generation
_1.3 Creating a sound clip
_1.4 Ethical and Social Impacts
_1.5 Past and Present of Generation Models
_1.6 How to develop a generative AI model
_1.7 Summary

Chapter 2 Transformers
_2.1 Language Model Utilization Cases
_2.2 Transformer Block
_2.3 Transformer Model Lineage
_2.4 The Power of Pre-Learning
_2.5 Transformers Summary
_2.6 Text generation project using language models
_2.7 Summary
_Practice Problems
_Challenge
_References

Chapter 3 Information Compression and Representation
_3.1 Autoencoder
_3.2 Variational Autoencoder
_3.3 CLIP
_3.4 Alternatives to CLIP
_3.5 Semantic-Based Image Retrieval Project
_3.6 Summary
_Practice Problems
_Challenge
_References

Chapter 4 Diffusion Model
_4.1 Core Principle: Iterative Refinement
_4.2 Diffusion Model Learning
_4.3 In-depth analysis of noise schedules
_4.4 In-Depth Analysis of U-Net and Alternatives
_4.5 In-depth analysis of diffusion goals
_4.6 Unconditional Diffusion Model Learning Project
_4.7 Summary
_Practice Problems
_Challenge
_References

Chapter 5 Stable Diffusion and Conditional Generation
_5.1 Adding Conditions for the Conditional Diffusion Model
_5.2 Potential Diffusion to Increase Efficiency
_5.3 In-Depth Analysis of Stable Diffusion Components
_5.4 Annotated sampling loop
_5.5 Open Data, Open Model
_5.6 Creating an Interactive Machine Learning Demo with Gradio Project
_5.7 Summary
_Practice Problems
_Challenge
_References

Transfer Learning for Generative Models Part 2

Chapter 6: Fine-Tuning Language Models
_6.1 Text Classification
_6.2 Text Generation
_6.3 Instructions
_6.4 Adapter Introduction
_6.5 Introduction to Quantization
_6.6 Integrated Implementation
_6.7 Deeper Understanding of Evaluation Methods
_6.8 Search Augmentation Generation Project
_6.9 Summary
_Practice Problems
_Challenge
_References

Chapter 7 Stable Diffusion Fine Tuning
_7.1 Stable Diffusion Full Model Fine Tuning
_7.2 Dream Booth
_7.3 LoRA Learning
_7.4 Adding new features to stable diffusion
_7.5 SDXL Dreambooth LoRA Learning Project
_7.6 Summary
_Practice Problems
_Challenge
_References

[Part 3: Moving on]

Chapter 8: Creative Use of Text-to-Image Models
_8.1 Image-to-Image Conversion
_8.2 Inpainting
_8.3 Editing prompt weights and images
_8.4 Editing real images with inversion
_8.5 ControlNet
_8.6 Image Prompting and Image Transformation
_8.7 Creative Drawing Creation Project
_8.8 Summary
_Practice Problems
_References

Chapter 9: Audio Generation
_9.1 Audio Data
_9.2 Speech-to-text conversion using a transformer-based architecture
_9.3 Text to speech, generated audio
_9.4 Audio Generation System Evaluation
_9.5 Future Development Direction
_9.6 End-to-End Dialogue System Project
_9.7 Summary
_Practice Problems
_Challenge
_References

Chapter 10: Advances and Recent Trends in Generative AI
_10.1 Preference Optimization
_10.2 Long Context
_10.3 Expert Mix
_10.4 Optimization and Quantization
_10.5 data
_10.6 A single model that solves everything
_10.7 Computer Vision
_10.8 3D Computer Vision
_10.9 Video Creation
_10.10 Multi-modality
_10.11 Community

APPENDIX A.
open source tools
APPENDIX B. LLM Memory Requirements
APPENDIX C.
End-to-end search augmentation generation
","
Detailed image
Detailed Image 1
","
Publisher's Review
A cutting-edge practical guide to generative AI, complete with hands-on exercises that design the inside and outside of models.

"Hands-On Generative AI" is a hands-on guide designed to help you understand the core principles of generative AI, implement them yourself, and apply them in practice.
We will explain key technologies such as transformers, autoencoders, and diffusion models, and learn the entire process of generating diverse multi-modal data such as text, images, and audio, along with actual code.

The first part of the book explains the concepts and basic structure of generative AI and introduces methods for generating text and images using pre-trained models.
The second half covers fine-tuning techniques based on transformers and diffusion models, RAG implementation, and practical examples utilizing cutting-edge technologies such as LoRA and Dreambooth, helping you gain a deeper understanding of the principles of generative AI.
The second half presents various strategies for applying generative AI in practice, with creative use cases such as audio generation, conditional image generation, and text-to-image applications.
We also cover the latest developments and trends in the field of generative AI, helping you grow into a responsible developer who utilizes generative AI technology.
Through this book, you will gain a solid understanding of generative AI, from its fundamentals to its practical applications, and further strengthen your development capabilities to keep pace with the latest technological trends.


Who is this book for?
● Developers who want to deeply understand the operating principles of generative AI models
● Those who want to go beyond simple API calls and directly fine-tune a model suitable for a specific domain
Researchers and engineers who want to build their own generative models using the Hugging Face ecosystem.
● LLM, those who have read the latest papers on diffusion models but are having difficulty implementing them in actual code
"]
GOODS SPECIFICS
- Date of issue: June 30, 2025
- Page count, weight, size: 472 pages | 1,086g | 183*235*24mm
- ISBN13: 9791169213981
- ISBN10: 1169213987

You may also like

카테고리