- عنوان کتاب: Building Resilient Distributed Systems
- نویسنده: Sam Newman
- حوزه: سیستمهای توزیعشده
- تعداد صفحه: 520
- زبان اصلی: انگلیسی
- نوع فایل: pdf
- حجم فایل: 7.74 مگابایت
سیستمهای توزیعشده در اشکال و ابعاد گوناگونی وجود دارند و بهطور فزایندهای به بخش مهمی از شیوه ساخت و ارائه نرمافزار تبدیل شدهاند. با این حال، سیستمهای توزیعشده چالشهایی را نیز به همراه میآورند. در نگاه نخست، به نظر میرسد که آنها امکان خلق نرمافزارهایی پایدارتر را فراهم میکنند؛ مگر نه اینکه ازکارافتادن یک ماشین واحد دیگر نباید کل سیستم ما را از دسترس خارج کند؟ متأسفانه، واقعیت به این سادگیها نیست. اگرچه سیستمهای توزیعشده در ظاهر ساده به نظر میرسند، اما پیچیدگی آنها از مجموعهای از تعاملاتِ در ابتدا ساده میان رایانهها ناشی میشود. ممکن است فراخوانیها بیش از حد طول بکشند یا با شکست مواجه شوند، تجهیزات شبکه آتش بگیرند، خطاهای پیکربندی بستههای شبکه شما را به بیراهههای اینترنت بفرستند، مراکز داده بر اثر درگیریهای منطقهای از کار بیفتند، هدف حمله کورکورانه و توزیعشدهی «محرومسازی از سرویس» (DDoS) قرار گیرید، یا حتی یک نقص ساده در حافظه پنهان (کش) کل سیستمتان را از کار بیندازد. اوضاع از این هم پیچیدهتر میشود؛ یک سیستم توزیعشده شامل افرادی است که آن را اداره و از آن استفاده میکنند، و همین افراد نیز میتوانند خود عاملی برای بیثباتی سیستم باشند. با این وجود، عامل انسانی حیاتیترین عنصر در مسیر تابآور ساختن سیستمهای توزیعشده ماست. در این کتاب، ماهیت سیستمهای توزیعشده و مفهوم کلی تابآوری را بررسی میکنم. در طول ۱۶ فصل پیشِ رو، شما را به سفری میبرم که در آن به همه چیز خواهیم پرداخت: از مفهوم ساده «مهلت زمانی» (Timeout) گرفته تا فروپاشی ساختمان؛ از سازوکار «تلاش مجدد» (Retry) تا «قضیه CAP»؛ و از «محدودسازی نرخ» (Rate Limiting) تا «سیستمهای اجتماعی-فنی». در پایان، درک بسیار عمیقتری از مفهوم تابآوری خواهید داشت و با راهکارهای افزایش تابآوری سیستم—هم از منظر فنی و هم از دیدگاه انسانی—بهخوبی آشنا خواهید شد.
Distributed systems come in many shapes and sizes, and they are increasingly a major part of how we build and deliver software. But distributed systems also create challenges. At face value, they appear to offer the ability to create more stable software. No longer should the death of a single machine take our system offline, right? Unfortunately, the world is not that simple. While distributed systems seem simple on the surface, complexity emerges from a set of initially straightforward interactions between computers. Calls can take too long or fail, networking equipment can catch fire, configuration errors can send your network packets down the wrong trouser leg of the internet, data centers can be taken out by regional conflict, you could be hit by an indiscriminate distributed denial of service attack, or a simple cache failure could bring your whole system down. It gets worse. A distributed system also consists of the people who operate and use the system, and they can be their own sources of instability. However, the human component is also the most vital one in taking our distributed systems and making them resilient. In this book, I explore the nature of distributed systems and the notion of resiliency in general. Through the following 16 chapters, I’ll take you on a journey in which we look at everything from the humble timeout to building collapse; retrying calls to CAP theorem; rate limiting to sociotechnical systems. By the end, you’ll know a lot more about how to make your system more resilient from both a technical and a people point of view, and you’ll have a much more rounded understanding of the notion of resilience in general.
این کتاب را میتوانید از لینک زیر بصورت رایگان دانلود کنید:
Download: Building Resilient Distributed Systems





نظرات کاربران