Previous
Build AI-Enhanced Web Apps: How to get reliable results with React, Next.js, and Vercel

Build AI-Enhanced Web Apps: How to get reliable results with...

$21.99
Next

Cherry Baby: A Novel

$7.99
Cherry Baby: A Novel

Building Multimodal Generative AI and Agentic Applications: Shaping...

Author: Indrajit Kar Language: EnglishPublisher: BPB PublicationsPages: 582Year: 2026
$ USD
  • $ USD
  • ₦ NGN
  • € EUR
  • £ GBP
  • $ CAD

$9.99

🔒 Secure payments powered by Paystack, a Stripe company
📥 Instant download after payment

Add to Wishlist
Add to Wishlist

Description

Building Multimodal Generative AI and Agentic Applications: Shaping concept to code for the future of multimodal and advanced agentic GenAI applications

Master the next generation of artificial intelligence by building multimodal and agentic AI systems capable of reasoning, generating content, and performing complex tasks across diverse data types.

Generative AI and agentic AI are transforming the way organizations interact with information. Modern AI systems can now understand text, images, speech, documents, and structured data while making intelligent decisions, coordinating multiple agents, and automating sophisticated workflows. As these technologies become central to enterprise software and research, the ability to design and deploy them has become an essential skill.

This comprehensive guide provides a practical roadmap for developing production-ready multimodal and agentic AI applications. Covering both core concepts and advanced implementation techniques, the book combines clear explanations with hands-on coding examples, architectural guidance, and real-world use cases to help readers build scalable AI solutions with confidence.

You’ll explore the technologies that power modern AI systems, including vision-language models, retrieval-augmented generation (RAG), vector databases, embeddings, intelligent agents, OCR, text-to-SQL, and hybrid AI architectures. Throughout the book, you’ll learn how to design efficient pipelines, orchestrate multiple AI agents, integrate human oversight, and deploy robust systems capable of operating in real-world production environments.

By the end of the book, you’ll have the knowledge and practical experience needed to create intelligent applications that retrieve information, generate high-quality responses, coordinate autonomous workflows, and scale reliably for enterprise use.

What You’ll Learn

  • Understand the principles behind multimodal generative AI and agentic AI systems.
  • Design architectures using retrieval-augmented generation (RAG), vector databases, embeddings, reranking, and agent planning.
  • Build efficient RAG pipelines for knowledge retrieval and content generation.
  • Develop human-in-the-loop workflows and coordinate multiple AI agents.
  • Create text-to-SQL systems for natural language database interaction.
  • Build OCR solutions to extract and process information from images and documents.
  • Combine traditional machine learning models with modern generative AI workflows.
  • Deploy, monitor, evaluate, and optimize production-grade AI systems using modern LLMOps practices.

Who This Book Is For

This book is designed for AI practitioners, data scientists, machine learning engineers, software developers, and technology professionals with a basic understanding of Python and machine learning. It is especially valuable for technical leads, enterprise architects, researchers, and anyone looking to build scalable, production-ready multimodal and agentic AI applications.

Topics Covered

  • Foundations of modern Generative AI
  • Multimodal AI architectures and design principles
  • Local and API-based GenAI implementations
  • Agentic AI systems with human oversight
  • Multi-stage AI workflows and orchestration
  • Bidirectional multimodal retrieval systems
  • Multimodal Retrieval-Augmented Generation (RAG)
  • Reranking and retrieval optimization techniques
  • Voice-enabled multimodal AI applications
  • Advanced multimodal system design and implementation
  • Natural language database querying with Text-to-SQL
  • Agent-driven Text-to-SQL architectures
  • Optical Character Recognition (OCR) for images and documents
  • Hybrid AI systems combining traditional ML and Generative AI
  • LLM operations, monitoring, evaluation, and production best practices

Whether you’re building enterprise AI platforms, intelligent assistants, automation pipelines, or next-generation multimodal applications, this guide provides the practical skills and architectural insight needed to transform cutting-edge AI concepts into real-world solutions.

Reviews

There are no reviews yet.

Be the first to review “Building Multimodal Generative AI and Agentic Applications: Shaping...”

Your email address will not be published. Required fields are marked *

Shopping cart

0
image/svg+xml

No products in the cart.

Continue Shopping