0

دانلود کتاب اصول زیرساخت پردازنده گرافیکی انویدیا، راهنمای ساختاریافته برای زیرساخت پردازنده گرافیکی انویدیا، از CUDA تا عملیات تولید

  • عنوان کتاب: NVIDIA GPU Infrastructure Fundamentals A structured guide to NVIDIA GPU infrastructure, from CUDA to production operations
  • نویسنده: Aranha, Vivian
  • حوزه: معماری کامپیوتر
  • سال انتشار: 2026
  • تعداد صفحه: 238
  • زبان اصلی: انگلیسی
  • نوع فایل: pdf
  • حجم فایل: 5.15 مگابایت

زیرساخت هوش مصنوعی دیگر تنها توسط GPU تعریف نمی‌شود. پلتفرم‌های شتاب‌دهنده مدرن، سخت‌افزار GPU را با نرم‌افزار سیستم، شبکه‌های پرسرعت، ذخیره‌سازی، تنظیم، نظارت، مجازی‌سازی، MLOps و فناوری‌های استنتاج ترکیب می‌کنند. هر لایه بخش متفاوتی از مشکل را حل می‌کند، اما چالش واقعی برای متخصصان زیرساخت، درک چگونگی تعامل این لایه‌ها و جایی است که مسئولیت یک فناوری به پایان می‌رسد و فناوری دیگر آغاز می‌شود. اکوسیستم NVIDIA این وسعت را منعکس می‌کند. CUDA پایه و اساس محاسبات شتاب‌دهنده GPU را فراهم می‌کند، در حالی که فناوری‌هایی مانند Tensor Cores، NVLink و NVSwitch بر نحوه اجرای و مقیاس‌بندی بارهای کاری تأثیر می‌گذارند. MIG و vGPU رویکردهای متفاوتی برای اشتراک‌گذاری منابع شتاب‌دهنده ارائه می‌دهند، DCGM از نظارت و مدیریت GPU پشتیبانی می‌کند و Kubernetes و Slurm مدل‌های برنامه‌ریزی متفاوتی ارائه می‌دهند. فراتر از خود GPU، فناوری‌هایی مانند InfiniBand، RDMA، GPUDirect Storage، BlueField DPUs و DOCA نحوه حرکت داده‌ها را در پلتفرم شکل می‌دهند. در لایه‌های کاربردی و عملیاتی، ابزارهایی مانند NGC، Airflow، MLflow، Kubeflow، ONNX، TensorRT و Triton به انتقال حجم کار هوش مصنوعی از توسعه به تولید کمک می‌کنند. با وجود فناوری‌های بسیار زیاد، یادگیری آنها به عنوان کاتالوگی از محصولات مجزا به ندرت کافی است. تصمیم در مورد زیرساخت معمولاً به حجم کار و مرزی که در آن یک مشکل رخ می‌دهد بستگی دارد. نیاز به منابع GPU ایزوله از سخت‌افزار با نیاز به ماشین‌های مجازی مجهز به GPU متفاوت است. مدلی که تشنه داده‌های ورودی است، به پاسخی متفاوت از مدلی که وابسته به محاسبات است، نیاز دارد. کانتینری که نمی‌تواند به GPU دسترسی داشته باشد، مشکلی متفاوت از یک پاد که قابل برنامه‌ریزی نیست یا یک سرویس استنتاج که توان عملیاتی ضعیفی دارد، ایجاد می‌کند. درک این تمایزات همان چیزی است که دانش محصول را به قضاوت مفید در مورد زیرساخت تبدیل می‌کند. این کتاب زیرساخت GPU NVIDIA را به عنوان یک سیستم متصل ارائه می‌دهد. قبل از معرفی CUDA و پشته نرم‌افزار NVIDIA، با مبانی حجم کار هوش مصنوعی و محاسبات شتاب‌یافته آغاز می‌شود. سپس معماری و ویژگی‌های GPUهای مرکز داده NVIDIA را بررسی خواهید کرد و یاد خواهید گرفت که چگونه الزامات حجم کار بر انتخاب پلتفرم تأثیر می‌گذارد. از آنجا، کتاب به لایه‌های عملیاتی پیرامون شتاب‌دهنده، از جمله پارتیشن‌بندی GPU، نظارت، برنامه‌ریزی، ذخیره‌سازی، شبکه‌سازی، مجازی‌سازی و تخلیه بار زیرساخت مبتنی بر DPU، می‌پردازد. فصل‌های بعدی این زیرساخت را به چرخه حیات هوش مصنوعی متصل می‌کنند. خواهید دید که چگونه Airflow، MLflow و Kubeflow از گردش‌های کاری یادگیری ماشینی قابل تکرار پشتیبانی می‌کنند؛ چگونه ONNX، TensorRT و Triton نقش‌های مختلفی را در استنتاج تولید ایفا می‌کنند؛ و چگونه NGC، Kubernetes، جعبه ابزار NVIDIA Container، اپراتور GPU و فناوری‌های نظارت در یک پلتفرم GPU قابل اجرا ترکیب می‌شوند. سناریوهای عیب‌یابی در سراسر کتاب، این ایده را تقویت می‌کنند که زیرساخت باید به عنوان مجموعه‌ای از لایه‌های متصل بررسی شود، نه به عنوان اجزای جداگانه. هدف این نیست که شما هر محصول یا مشخصات NVIDIA را به خاطر بسپارید. نسل‌های GPU، نسخه‌های نرم‌افزاری و قابلیت‌های پلتفرم همچنان در حال تکامل خواهند بود. در عوض، این کتاب بر دانش پایدارتر تمرکز دارد: اینکه هر فناوری برای انجام چه کاری طراحی شده است، چگونه با پشته اطراف خود ارتباط دارد، چه بده‌بستان‌هایی فناوری‌های مشابه را متمایز می‌کند و چگونه می‌توان از یک نیاز بار کاری یا علامت عملیاتی به یک پاسخ زیرساختی مناسب استدلال کرد. در پایان کتاب، شما باید یک مدل ذهنی ساختاریافته داشته باشید که به شما امکان می‌دهد با اعتماد به نفس بیشتری در مورد زیرساخت‌های پردازنده‌های گرافیکی انویدیا بحث، ارزیابی و کار کنید. این کتاب برای مدیران سیستم، متخصصان فضای ابری و DevOps، تیم‌های مرکز داده و شبکه، معماران راهکار، مدیران فنی، متخصصان پیش‌فروش و هر کسی که به سمت نقش‌های زیرساخت هوش مصنوعی می‌رود و نیاز به درک ساختاریافته‌ای از اکوسیستم پردازنده‌های گرافیکی انویدیا دارد، مناسب است. همچنین برای خوانندگانی که در کنار تیم‌های هوش مصنوعی، MLOps یا پلتفرم کار می‌کنند و می‌خواهند درک کنند که چگونه محاسبات، شبکه، ذخیره‌سازی، هماهنگی و استنتاج با هم هماهنگ می‌شوند، مناسب است. آشنایی اولیه با مفاهیم فناوری اطلاعات، فضای ابری، شبکه یا مرکز داده مفید خواهد بود، اما هیچ تجربه قبلی با پردازنده‌های گرافیکی انویدیا لازم نیست. برای دنبال کردن کتاب نیازی به پیش‌زمینه برنامه‌نویسی یا علوم داده ندارید. تمرکز بر مفاهیم زیرساخت، مرزهای فناوری، تصمیمات عملیاتی و بده‌بستان‌های پلتفرم است.

AI infrastructure is no longer defined by the GPU alone. Modern accelerated platforms combine GPU hardware with system software, high-speed networking, storage, orchestration, monitoring, virtualization, MLOps, and inference technologies. Each layer solves a different part of the problem, but the real challenge for infrastructure professionals is understanding how those layers interact and where the responsibility of one technology ends and another begins. The NVIDIA ecosystem reflects this breadth. CUDA provides the foundation for GPUaccelerated computing, while technologies such as Tensor Cores, NVLink, and NVSwitch influence how workloads execute and scale. MIG and vGPU provide different approaches to sharing accelerator resources, DCGM supports GPU monitoring and management, and Kubernetes and Slurm provide different scheduling models. Beyond the GPU itself, technologies such as InfiniBand, RDMA, GPUDirect Storage, BlueField DPUs, and DOCA shape how data moves through the platform. At the application and operations layers, tools such as NGC, Airflow, MLflow, Kubeflow, ONNX, TensorRT, and Triton help move AI workloads from development into production. With so many technologies involved, learning them as a catalogue of individual products is rarely enough. An infrastructure decision usually depends on the workload and the boundary at which a problem occurs. A requirement for hardware-isolated GPU resources is different from a requirement for GPU-enabled virtual machines. A model that is starved for input data needs a different response from one that is compute-bound. A container that cannot access a GPU presents a different problem from a pod that cannot be scheduled or an inference service that has poor throughput. Understanding these distinctions is what turns product knowledge into useful infrastructure judgment. This book presents NVIDIA GPU infrastructure as one connected system. It begins with the foundations of AI workloads and accelerated computing before introducing CUDA and the NVIDIA software stack. You will then examine the architecture and characteristics of NVIDIA data center GPUs and learn how workload requirements influence platform selection. From there, the book moves into the operational layers surrounding the accelerator, including GPU partitioning, monitoring, scheduling, storage, networking, virtualization, and DPU-based infrastructure offload. The later chapters connect this infrastructure to the AI lifecycle. You will see how Airflow, MLflow, and Kubeflow support reproducible machine learning workflows; how ONNX, TensorRT, and Triton serve different roles in production inference; and how NGC, Kubernetes, the NVIDIA Container Toolkit, GPU Operator, and monitoring technologies combine into an operable GPU platform. Troubleshooting scenarios throughout the book reinforce the idea that infrastructure should be investigated as a set of connected layers rather than as isolated components. The aim is not to make you memorize every NVIDIA product or specification. GPU generations, software versions, and platform capabilities will continue to evolve. Instead, this book focuses on the more durable knowledge: what each technology is designed to do, how it relates to the surrounding stack, what trade-offs distinguish similar technologies, and how to reason from a workload requirement or operational symptom to an appropriate infrastructure response. By the end of the book, you should have a structured mental model that allows you to discuss, evaluate, and work with NVIDIA GPU infrastructure with greater confidence. This book is for system administrators, cloud and DevOps professionals, data center and networking teams, solution architects, technical managers, presales professionals, and anyone moving into AI infrastructure roles who needs a structured understanding of the NVIDIA GPU ecosystem. It is also suitable for readers who work alongside AI, MLOps, or platform teams and want to understand how compute, networking, storage, orchestration, and inference fit together. Basic familiarity with IT, cloud, networking, or data center concepts will be helpful, but no previous experience with NVIDIA GPUs is required. You do not need a programming or data science background to follow the book; the focus is on infrastructure concepts, technology boundaries, operational decisions, and platform trade-offs.

این کتاب را میتوانید از لینک زیر بصورت رایگان دانلود کنید:

Download: NVIDIA GPU Infrastructure Fundamentals

نظرات کاربران

  •  چنانچه دیدگاه شما توهین آمیز باشد تایید نخواهد شد.
  •  چنانچه دیدگاه شما جنبه تبلیغاتی داشته باشد تایید نخواهد شد.

دیدگاهتان را بنویسید

نشانی ایمیل شما منتشر نخواهد شد. بخش‌های موردنیاز علامت‌گذاری شده‌اند *

بیشتر بخوانید