- عنوان کتاب: NVIDIA GPU Infrastructure Fundamentals A structured guide to NVIDIA GPU infrastructure, from CUDA to production operations
- نویسنده: Aranha, Vivian
- حوزه: معماری کامپیوتر
- سال انتشار: 2026
- تعداد صفحه: 238
- زبان اصلی: انگلیسی
- نوع فایل: pdf
- حجم فایل: 5.15 مگابایت
زیرساخت هوش مصنوعی دیگر تنها توسط GPU تعریف نمیشود. پلتفرمهای شتابدهنده مدرن، سختافزار GPU را با نرمافزار سیستم، شبکههای پرسرعت، ذخیرهسازی، تنظیم، نظارت، مجازیسازی، MLOps و فناوریهای استنتاج ترکیب میکنند. هر لایه بخش متفاوتی از مشکل را حل میکند، اما چالش واقعی برای متخصصان زیرساخت، درک چگونگی تعامل این لایهها و جایی است که مسئولیت یک فناوری به پایان میرسد و فناوری دیگر آغاز میشود. اکوسیستم NVIDIA این وسعت را منعکس میکند. CUDA پایه و اساس محاسبات شتابدهنده GPU را فراهم میکند، در حالی که فناوریهایی مانند Tensor Cores، NVLink و NVSwitch بر نحوه اجرای و مقیاسبندی بارهای کاری تأثیر میگذارند. MIG و vGPU رویکردهای متفاوتی برای اشتراکگذاری منابع شتابدهنده ارائه میدهند، DCGM از نظارت و مدیریت GPU پشتیبانی میکند و Kubernetes و Slurm مدلهای برنامهریزی متفاوتی ارائه میدهند. فراتر از خود GPU، فناوریهایی مانند InfiniBand، RDMA، GPUDirect Storage، BlueField DPUs و DOCA نحوه حرکت دادهها را در پلتفرم شکل میدهند. در لایههای کاربردی و عملیاتی، ابزارهایی مانند NGC، Airflow، MLflow، Kubeflow، ONNX، TensorRT و Triton به انتقال حجم کار هوش مصنوعی از توسعه به تولید کمک میکنند. با وجود فناوریهای بسیار زیاد، یادگیری آنها به عنوان کاتالوگی از محصولات مجزا به ندرت کافی است. تصمیم در مورد زیرساخت معمولاً به حجم کار و مرزی که در آن یک مشکل رخ میدهد بستگی دارد. نیاز به منابع GPU ایزوله از سختافزار با نیاز به ماشینهای مجازی مجهز به GPU متفاوت است. مدلی که تشنه دادههای ورودی است، به پاسخی متفاوت از مدلی که وابسته به محاسبات است، نیاز دارد. کانتینری که نمیتواند به GPU دسترسی داشته باشد، مشکلی متفاوت از یک پاد که قابل برنامهریزی نیست یا یک سرویس استنتاج که توان عملیاتی ضعیفی دارد، ایجاد میکند. درک این تمایزات همان چیزی است که دانش محصول را به قضاوت مفید در مورد زیرساخت تبدیل میکند. این کتاب زیرساخت GPU NVIDIA را به عنوان یک سیستم متصل ارائه میدهد. قبل از معرفی CUDA و پشته نرمافزار NVIDIA، با مبانی حجم کار هوش مصنوعی و محاسبات شتابیافته آغاز میشود. سپس معماری و ویژگیهای GPUهای مرکز داده NVIDIA را بررسی خواهید کرد و یاد خواهید گرفت که چگونه الزامات حجم کار بر انتخاب پلتفرم تأثیر میگذارد. از آنجا، کتاب به لایههای عملیاتی پیرامون شتابدهنده، از جمله پارتیشنبندی GPU، نظارت، برنامهریزی، ذخیرهسازی، شبکهسازی، مجازیسازی و تخلیه بار زیرساخت مبتنی بر DPU، میپردازد. فصلهای بعدی این زیرساخت را به چرخه حیات هوش مصنوعی متصل میکنند. خواهید دید که چگونه Airflow، MLflow و Kubeflow از گردشهای کاری یادگیری ماشینی قابل تکرار پشتیبانی میکنند؛ چگونه ONNX، TensorRT و Triton نقشهای مختلفی را در استنتاج تولید ایفا میکنند؛ و چگونه NGC، Kubernetes، جعبه ابزار NVIDIA Container، اپراتور GPU و فناوریهای نظارت در یک پلتفرم GPU قابل اجرا ترکیب میشوند. سناریوهای عیبیابی در سراسر کتاب، این ایده را تقویت میکنند که زیرساخت باید به عنوان مجموعهای از لایههای متصل بررسی شود، نه به عنوان اجزای جداگانه. هدف این نیست که شما هر محصول یا مشخصات NVIDIA را به خاطر بسپارید. نسلهای GPU، نسخههای نرمافزاری و قابلیتهای پلتفرم همچنان در حال تکامل خواهند بود. در عوض، این کتاب بر دانش پایدارتر تمرکز دارد: اینکه هر فناوری برای انجام چه کاری طراحی شده است، چگونه با پشته اطراف خود ارتباط دارد، چه بدهبستانهایی فناوریهای مشابه را متمایز میکند و چگونه میتوان از یک نیاز بار کاری یا علامت عملیاتی به یک پاسخ زیرساختی مناسب استدلال کرد. در پایان کتاب، شما باید یک مدل ذهنی ساختاریافته داشته باشید که به شما امکان میدهد با اعتماد به نفس بیشتری در مورد زیرساختهای پردازندههای گرافیکی انویدیا بحث، ارزیابی و کار کنید. این کتاب برای مدیران سیستم، متخصصان فضای ابری و DevOps، تیمهای مرکز داده و شبکه، معماران راهکار، مدیران فنی، متخصصان پیشفروش و هر کسی که به سمت نقشهای زیرساخت هوش مصنوعی میرود و نیاز به درک ساختاریافتهای از اکوسیستم پردازندههای گرافیکی انویدیا دارد، مناسب است. همچنین برای خوانندگانی که در کنار تیمهای هوش مصنوعی، MLOps یا پلتفرم کار میکنند و میخواهند درک کنند که چگونه محاسبات، شبکه، ذخیرهسازی، هماهنگی و استنتاج با هم هماهنگ میشوند، مناسب است. آشنایی اولیه با مفاهیم فناوری اطلاعات، فضای ابری، شبکه یا مرکز داده مفید خواهد بود، اما هیچ تجربه قبلی با پردازندههای گرافیکی انویدیا لازم نیست. برای دنبال کردن کتاب نیازی به پیشزمینه برنامهنویسی یا علوم داده ندارید. تمرکز بر مفاهیم زیرساخت، مرزهای فناوری، تصمیمات عملیاتی و بدهبستانهای پلتفرم است.
AI infrastructure is no longer defined by the GPU alone. Modern accelerated platforms combine GPU hardware with system software, high-speed networking, storage, orchestration, monitoring, virtualization, MLOps, and inference technologies. Each layer solves a different part of the problem, but the real challenge for infrastructure professionals is understanding how those layers interact and where the responsibility of one technology ends and another begins. The NVIDIA ecosystem reflects this breadth. CUDA provides the foundation for GPUaccelerated computing, while technologies such as Tensor Cores, NVLink, and NVSwitch influence how workloads execute and scale. MIG and vGPU provide different approaches to sharing accelerator resources, DCGM supports GPU monitoring and management, and Kubernetes and Slurm provide different scheduling models. Beyond the GPU itself, technologies such as InfiniBand, RDMA, GPUDirect Storage, BlueField DPUs, and DOCA shape how data moves through the platform. At the application and operations layers, tools such as NGC, Airflow, MLflow, Kubeflow, ONNX, TensorRT, and Triton help move AI workloads from development into production. With so many technologies involved, learning them as a catalogue of individual products is rarely enough. An infrastructure decision usually depends on the workload and the boundary at which a problem occurs. A requirement for hardware-isolated GPU resources is different from a requirement for GPU-enabled virtual machines. A model that is starved for input data needs a different response from one that is compute-bound. A container that cannot access a GPU presents a different problem from a pod that cannot be scheduled or an inference service that has poor throughput. Understanding these distinctions is what turns product knowledge into useful infrastructure judgment. This book presents NVIDIA GPU infrastructure as one connected system. It begins with the foundations of AI workloads and accelerated computing before introducing CUDA and the NVIDIA software stack. You will then examine the architecture and characteristics of NVIDIA data center GPUs and learn how workload requirements influence platform selection. From there, the book moves into the operational layers surrounding the accelerator, including GPU partitioning, monitoring, scheduling, storage, networking, virtualization, and DPU-based infrastructure offload. The later chapters connect this infrastructure to the AI lifecycle. You will see how Airflow, MLflow, and Kubeflow support reproducible machine learning workflows; how ONNX, TensorRT, and Triton serve different roles in production inference; and how NGC, Kubernetes, the NVIDIA Container Toolkit, GPU Operator, and monitoring technologies combine into an operable GPU platform. Troubleshooting scenarios throughout the book reinforce the idea that infrastructure should be investigated as a set of connected layers rather than as isolated components. The aim is not to make you memorize every NVIDIA product or specification. GPU generations, software versions, and platform capabilities will continue to evolve. Instead, this book focuses on the more durable knowledge: what each technology is designed to do, how it relates to the surrounding stack, what trade-offs distinguish similar technologies, and how to reason from a workload requirement or operational symptom to an appropriate infrastructure response. By the end of the book, you should have a structured mental model that allows you to discuss, evaluate, and work with NVIDIA GPU infrastructure with greater confidence. This book is for system administrators, cloud and DevOps professionals, data center and networking teams, solution architects, technical managers, presales professionals, and anyone moving into AI infrastructure roles who needs a structured understanding of the NVIDIA GPU ecosystem. It is also suitable for readers who work alongside AI, MLOps, or platform teams and want to understand how compute, networking, storage, orchestration, and inference fit together. Basic familiarity with IT, cloud, networking, or data center concepts will be helpful, but no previous experience with NVIDIA GPUs is required. You do not need a programming or data science background to follow the book; the focus is on infrastructure concepts, technology boundaries, operational decisions, and platform trade-offs.
این کتاب را میتوانید از لینک زیر بصورت رایگان دانلود کنید:
Download: NVIDIA GPU Infrastructure Fundamentals





نظرات کاربران