- عنوان کتاب: Machine Learning Systems at Scale , Vol 2
- نویسنده: Vijay Janapa Reddi
- حوزه: یادگیری ماشین
- تعداد صفحه: 1170
- زبان اصلی: انگلیسی
- نوع فایل: pdf
- حجم فایل: 19.5 مگابایت
این کتاب برای هر کسی است که یادگیری ماشینی را روی یک ماشین واحد درک میکند و اکنون با چالش کار کردن آن در مقیاس بزرگ مواجه است. مشکلات مقیاس، مشکلات یک گره واحد نیستند، بلکه بزرگترند. آنها از نظر کیفی متفاوت هستند: تقسیمبندی شبکهها، سختافزار به عنوان یک قطعیت آماری شکست میخورد و تأثیر اجتماعی هر تصمیم طراحی توسط میلیونها کاربر تقویت میشود. دامنه این کتاب از طراحی بستر فیزیکی یک مرکز داده هوش مصنوعی گرفته تا مقیاسبندی آموزش فراتر از محدودیتهای یک شتابدهنده واحد، و استدلال در مورد آنچه هنگام خدمترسانی یک ناوگان تولیدی به یک پایگاه کاربر جهانی اتفاق میافتد، متغیر است. در کل، تمرکز بر فیزیک توزیع است – متغیرهایی که هر سیستم یادگیری ماشینی در مقیاس بزرگ را صرف نظر از چارچوب، مدل یا دوره زمانی کنترل میکنند. چرا یک کتاب درسی در سطح ناوگان؟ در سال ۲۰۱۲، آموزش AlexNet پنج تا شش روز روی دو پردازنده گرافیکی (GPU) طول کشید (Krizhevsky و همکاران، ۲۰۱۲). تا سال ۲۰۲۳، تخمینهای عمومی کلاس GPT-4، آموزش مرزی را تقریباً ۲۵۰۰۰ پردازنده گرافیکی در حال اجرا به مدت حدود سه ماه قرار داد (SemiAnalysis 2023)؛ OpenAI پیکربندی واقعی سختافزار را فاش نکرد (OpenAI و همکاران، ۲۰۲۳). این همان مشکل مهندسی در مقیاس بزرگتر نیست. این یک مشکل مهندسی کاملاً متفاوت است. در یک ماشین واحد، عملکرد توسط دیوار حافظه – شکاف بین سرعت پردازنده و پهنای باند حافظه – اداره میشود. در یک خوشه توزیعشده، یک دیوار جدید پدیدار میشود: دیوار پهنای باند دوبخشی، که در آن حداقل پهنای باند شبکهای که خوشه را به دو بخش تقسیم میکند، کل توان عملیاتی سیستم را محدود میکند. دادهها دیگر تنها از طریق سلسله مراتب محلی حرکت نمیکنند؛ بلکه از پارچههای نوری که توسط سرعت نور بین رکها و قارهها کنترل میشوند، عبور میکنند. خرابیهای سختافزاری، که به اندازه کافی در یک ماشین واحد نادر هستند که به عنوان استثنا در نظر گرفته شوند، در مقیاس ناوگان به رویدادهای آماری معمول تبدیل میشوند. یک کار آموزشی که برای خرابی برنامهریزی نکرده باشد، با شکست مواجه خواهد شد. اینها مزاحمتهای عملیاتی نیستند. آنها محدودیتهای فیزیکی هستند – به همان اندازه اساسی و دائمی که دیوار حافظه، قانون آمدال یا هزینه انرژی جابجایی دادهها. آنها به رشته مهندسی خاص خود، با ثابتهای خاص خود نیاز دارند: • قضیه پهنای باند دوبخشی: عملکرد یک سیستم توزیعشده توسط حداقل پهنای باندی که توپولوژی شبکه را به دو قسمت تقسیم میکند، محدود میشود. • قانون ناوگان: هر مرحله آموزشی، مالیات ارتباطی پرداخت میکند که با اندازه خوشه افزایش مییابد در حالی که محاسبات هر گره کاهش مییابد. • قانون ایست بازرسی یانگ-دالی: فاصله ایست بازرسی بهینه، هزینه نوشتن حالت را در مقابل هزینه مورد انتظار کار از دست رفته به دلیل خرابی متعادل میکند. • قانون تسلط بر هزینه خدمت: برای مدلهای تولید پرترافیک، هزینههای عملیاتی استنتاج تجمعی میتواند از هزینه آموزش یکباره فراتر رود و کارایی خدمت را به یک هدف بهینهسازی چرخه عمر تبدیل کند. • قانون عدم امکان انصاف: هیچ سیستمی نمیتواند همزمان کالیبراسیون، شانسهای برابر و برابری جمعیتی را برآورده کند، زمانی که نرخهای پایه بین گروهها متفاوت است. اگر مهندسی سیستمهای یادگیری ماشین تک ماشینی بپرسد “یک ماشین چه چیزی میتواند بسازد؟”، مهندسی سیستمهای یادگیری ماشین در مقیاس ناوگانی میپرسد “هزار ماشین چه چیزی میتوانند با هم بسازند و با چه هزینهای برای قابلیت اطمینان، کارایی و جامعه؟”
This book is for anyone who understands machine learning on a single machine and now faces the challenge of making it work at scale. The problems of scale are not the problems of a single node, only bigger. They are qualitatively different: networks partition, hardware fails as a statistical certainty, and the societal impact of every design decision is amplified by millions of users. The scope ranges from designing the physical substrate of an AI data center, to scaling training beyond the limits of a single accelerator, to reasoning about what happens when a production fleet serves a global user base. Throughout, the focus is on the physics of distribution—the invariants that govern every fleet-scale ML system regardless of the framework, the model, or the era. Why a Fleet-Level Textbook In 2012, training AlexNet took five to six days on two GPUs (Krizhevsky et al. 2012). By 2023, public GPT-4-class estimates placed frontier training at roughly 25,000 GPUs running for about three months (SemiAnalysis 2023); OpenAI did not disclose the actual hardware configuration (OpenAI et al. 2023). This is not the same engineering problem at a larger scale. It is a categorically different engineering problem. On a single machine, performance is governed by the memory wall—the gap between processor speed and memory bandwidth. In a distributed cluster, a new wall emerges: the bisection bandwidth wall, where the minimum network bandwidth that bisects the cluster caps total system throughput. Data no longer moves through local hierarchies alone; it traverses optical fabrics governed by the speed of light between racks and across continents. Hardware failures, rare enough on a single machine to be treated as exceptions, become routine statistical events at fleet scale. A training job that does not plan for failure will fail. These are not operational annoyances. They are physical constraints—as fundamental and as permanent as the memory wall, Amdahl’s Law, or the energy cost of data movement. They require their own engineering discipline, with its own invariants: • The Bisection Bandwidth Theorem: The performance of a distributed system is limited by the minimum bandwidth that bisects the network topology. • The Fleet Law: Every training step pays a communication tax that grows with cluster size while per-node computation shrinks. • The Young-Daly Checkpoint Law: The optimal checkpoint interval balances the cost of writing state against the expected cost of lost work due to failure. • The Serving Cost Dominance Law: For high-traffic production models, cumulative inference operational expenditure can exceed the one-time training cost, making serving efficiency a lifecycle optimization target. • The Fairness Impossibility Law: No system can simultaneously satisfy calibration, equalized odds, and demographic parity when base rates differ between groups. If single-machine ML systems engineering asks “what can one machine build?”, then fleet-scale ML systems engineering asks “what can a thousand machines build together, and at what cost to reliability, efficiency, and society?”
این کتاب را میتوانید از لینک زیر بصورت رایگان دانلود کنید:





نظرات کاربران