# Reliability is a Question — transcript source
## 公式アブストラクト
One thing that we all share—big environments, small environments, and everything in between, no matter what your organization does for a living—is a thirst for reliability. We spend (or perhaps should spend) large swaths of our time asking our systems to provide the appropriate level of reliability. And when they don’t, we then try to ask why. Sometimes we even get the hint of an answer.
The problem is that the obvious questions don’t always lead to the right answers, relationships can be counter-intuitive, and learning from failure is not easy even if, or perhaps especially if, you approach things using a Site Reliability Engineering (SRE) mindset. Let’s talk about these questions and how they can lead you astray no matter how experienced you are with SRE. We will get you back on to the right path with better questions (and maybe even an answer or two!). You will be taking away some concrete approaches to your challenges around failures, communication, organizational structure, monitoring, and a few other surprise topics. Come join us for a chance to question your answers and answer your questions around reliability!
## 登壇者略歴
the editor/curator of Seeking SRE, and the author of Becoming SRE
David Blank-Edelman
David Blank-Edelman is the editor/curator of Seeking SRE: Conversations About Running Production Systems at Scale (O’Reilly) and the author of Becoming SRE: First Steps Toward Reliability for You and Your Organization (O’Reilly). David is a co-founder of the SREcon conference and has roughly 40 years of experience in the operations space.
## タイムスタンプ付き逐語文字起こし
**[00:46]**
**司会:** David Blank-Edelman
**[00:57]**
**David Blank-Edelman:** Hello, my dear Japanese friends. I'm so honored to be here. I want to thank Yohei and Mary for all they did to bring this conference into existence. Please thank them if you would.
(拍手)
So as you've probably figured out by now, I don't know very much Japanese, and you wouldn't want to hear the Japanese I know. So I appreciate you being willing to hear this in translation. And I'm hoping we'll have a chance, all of us, to talk together later. I have good translation technology and would love to chat with you afterwards.
So, tonight—or today—we're going to be talking about questions. And in fact, what I'm trying to do as part of being inclusive is I'm going to be looking at some of the questions that some people ask about reliability, that people have asked me, and teach you how to ask better questions. We're going to take the questions apart. We're going to figure out how they could be better. We're going to figure out what questions we should ask instead. And that will hopefully get you to a much deeper place when it comes to SRE. And if you've been doing SRE for a while, I think you'll find some of these questions sound very simple, but you know from your experience that they're not as simple as that.
And just in case you're concerned, I will indeed be playing—paying the AI tax. That's the thing that says that I'm required to talk about AI at some time during this talk. And I promise you that I will get to AI later, towards the end.
Okay, so let's begin. I thought that of any of the people in the world who would appreciate starting with some poetry, it would be the Japanese folks that I have met. And so I want to show you this, and then I'm going to read it in English because I don't know this Japanese, I don't know how easy it will be for you to read from this distance, but perhaps you'll be able to hear it in translation. I don't trust the poetry of the AI translation.
I should also say that I get very excited when I talk, which means I'm probably going to talk quickly, which means if you see that the AI translation is starting to have smoke coming out of it, you should make the international sign for 'slow down' and I will do my best to slow down. Hopefully that makes sense.
So this is—this is a letter written by the poet Rilke, a German poet. And I think it's a good way to start. Let me read to you this in English because it's the only one—only way I can do it.
So, 'I would like to beg of you, dear sir, as well as I can, to have patience with everything unresolved in your heart and to try to love the questions themselves as if they were locked rooms or books written in a very foreign language. Don't search for the answers, which could not be given to you now because you would not be able to live them. And the point is to live everything. Live the questions now. Perhaps then, someday, far in the future, you will gradually, without even noticing it, live your way into the answers.' And so this talk is, again, about questions. There might not be as many answers as you'd like, but if you learn how to make better questions for reliability, I think it will help you.
So let's start off with the first one, which is this one right here: 'Is my something (system, service, etc.) reliable?'
And this is my way of trying to talk to you about expanding your idea of what reliability is, because most of the time when people think about reliability, they think about availability, right? Is the system or the service up, and is it down? And in fact, as an SRE, you probably will spend most of your time worrying about whether the system is up or if it's down. But it's not the only expectation that we have when we talk about reliability.
For example, we talk about latency, right? Latency, how quickly do we get a response back when we make a request? This is important in certain fields. For example, at Microsoft, at Xbox, where I used—you know, I used to work at Microsoft—at Xbox they care a lot about latency. And there's a saying—and I don't know how well it translates to Japanese—which is 'slow is the new down' because you know and I know that if your service is too slow and responds too slowly, people will go somewhere else. So that's why latency is important.
But maybe you're dealing with pipelines and you care about throughput. Or maybe you're dealing with batch processing and you care about coverage—did I process all my data? Or maybe you're thinking a little bit about correctness, which is something we don't think a lot about—did my code do what I wanted it to do? And am I measuring that?
Then there's fidelity. Fidelity is a little harder to talk about because it's meant to indicate how often you give the full experience to the customer. So, for example, if you watch Netflix and you go to their web page, on their web page they have different sections, right? They have recommendation engines, they have what they're playing now, they have what you watched before. And if the recommendation engine isn't working, you don't get a white screen that says 'Sorry, no movies today.' Instead, what they do is they take that—what that part out of the—out of the page, and put something static there. And so fidelity is asking how often do we deliver the experience we expected to the customer.
If you deal with elections or sports scores, maybe you care about freshness. And if anybody in this room deals with storage, you probably care a lot when you write a bit, that you can read that bit again later.
Now, these are many facets of the same thing. And one thing that's really important about this is to understand—since we want to understand how to think about SRE—is that reliability is measured from the customer's perspective, not from the component perspective. I'll say it one more time just in case we didn't get it in translation: reliability is measured from the customer's perspective, not from the component perspective.
I would like to do a quiz now. Are the people in this room ready and okay with the quiz? Yes? Yes, no?
Remember, I can't tell how long it takes for the translation to happen, so I am working with unknown latency here as to how long. So, please be nice to me and cheer.
Okay, let's assume that whatever you do for a living right now, you decided to quit your job and you want to join a company that makes tote bags with clouds on them, because you hear cloud is going to be a really big thing someday, and you would like to start making money from that. In order to do that, you set up 100 servers.
And yes, this is 100, and yes, putting 100 of anything on a PowerPoint slide is a big pain, but I wanted to make sure you had the full experience.
Now, imagine that something happens at the data center, and there's a power surge, or maybe there's a bad update, and 14 of these machines burst into flame. Okay? So now we have a situation where there are 86 machines working just fine, and 14 of them are broken.
Are you ready for your quiz? I don't want to wait on just you to say this, because this is not just for him. Okay, here's your quiz. Is this situation:
A: No big deal, if you happen to be at the beach, drinking a fruity drink in a coconut with a paper umbrella—I don't know if this is going to translate—you can stay there and eventually go fix the problem.
Or is the situation:
B: You should walk to your desk right now and go work fix it.
Or is the situation:
C: It's a crisis, and if it's 2 AM in the morning, I'm still going to wake up the CEO and the CTO and all the people that are in charge.
Okay, now we're going to vote. Slowly.
How many people here in the room today think it's A? Please raise your hand for me. Okay, a few.
How many people in this room believe the answer is B? Okay, much more.
How many people in the room think the answer is C? Okay, a little less.
Would you like to know what I think the answer is? Yes?
For those of you who can't read English or can't see it from this distance, my shirt says 'It depends.' And the reason why it says 'It depends' is if none of your customers notice a problem when this happens, maybe it's A. If things are slower and it's a problem for your customers, maybe it's B. And if in fact, the thing that has gone down is the thing that's making your company or organization money or the thing you care about, it could very well be C. And so what I wanted to say is that your monitoring system would have told you '86 up, 14 down,' but that's not what matters.
Okay, second question.
So, the question that people ask often in different fields is: 'How do I get rid of all of my failures or my errors in my system?'
And to understand why I think this is a problematic question, I want to talk about what the SRE mindset is, the way SREs think, at least the way I think we think.
So, to my mind, the SRE mindset starts with curiosity. It is all about curiosity. And so specifically, often you'll see us ask the questions: 'How does the system work?' And you'll also hear us ask the question: 'How does the system fail?' And the reason why we ask the question 'How does the system fail?' is because we want to know 'How does the system work?' And so it's about this curiosity.
And in many ways, I think the mindset is about this question.
So, I also want to say when we talk about these things, we're asking 'How do these things work in production?' We're not asking 'How do things work on your whiteboard?' We're not asking 'How things work in your documentation'—if you have documentation. We're asking 'How do they really work?' 'How do they work for our customers?' 'How do they work in production?' 'How do they work for more people?' 'How do they work better?' 'How does it scale?' These are all questions that we are keenly interested in.
And so it's important to think about this because as we think about how to define this, again, for the customer, we want to think a little bit about our relationships. And our relationships that define SRE work are two that I think are a little strange. One of them is we have this relationship to people. Reliability work is relentlessly collaborative. I really hope that translates. We must collaborate to work on reliability work. You cannot do it in your cubicle alone, right? You have to bring the right people into the picture, you need to ask people, you need to talk to them.
The other thing that I think is counter-intuitive is that we have a different understanding and relationship to errors and failure than most of the world does. Most of the world would like to get rid of all the errors and all the failures. We, on the other hand, think about errors and failures as signal. We're looking for more information, again, 'How does the system work?' and 'How does it fail?' and these failures are the things that tell us. We don't have an adversarial relationship to this. Now, I'm not saying that we're thrilled when everything goes down, or we want more outages. I'm simply saying that we treat this as a way to learn from failure in a way that other places don't, who are just like, 'Yeah, I want to make that stop.'
So, when we talk about 'How do I get rid of all my errors and failures?', we're really talking—we're really—that's not the right question, I guess is what I want to say. You know, the question is more 'How do I learn from them?' Now, ideally what you're looking for are novel failures and novel errors, or ephemeral ones, ones that go away. If you're always having the same thing go down every single time, that's not good. You know what the problem is, you just have to fix it.
Okay, question number three.
So, if you'd like to see me get really passionate and angry, we can have a conversation about 'root cause' and the words 'root cause' here.
Please tell me you don't use this in Japanese. Is that true? Or is root cause something we talk about here as well a lot? Yes, no? I need—I need some information. Yes? Okay, great. So I'm going to tell you now to stop. And I'm going to explain why I'm telling you to stop do that, because I'm going to tell you a little story. Okay? Here's—here's just a small story. Now, remember when we were discussing that company that sold tote bags? Now imagine that Pat walks into the data center and trips over a cable that was laid by Oscar, who is installing the server. Now imagine that in that process, the database that was running on that server goes down. This database was configured by Susan. And in fact, how the data on that database and all your databases were—were handled by Yasmin. Now, Niraj wrote the back-end. And the back-end's having some problems right now since we pulled the plug on this database, and it's starting to get slower and slower, which is a real problem for the front-end, which is getting even worse. Liz's monitoring picks this up, but picks it up a little late, and in fact sends a message out, and people don't get the message in time. And then Sam, who has configured the load balancer—that load balancer starts adding—sending—handing out 503s like candy. Hope that makes sense, do you know what a 503 is? Yes. Yes?
So, my question for you is: somebody is about to go buy your product. Who is responsible for the outage and the loss of that sale? Which person, please?
I really hope you look at this and you say 'None of them and all of them—it's the system.'
And so when we talk about root cause things, we're usually setting up a framework or framing that there is one causal problem that causes this thing. And because we're saying there's a root cause, we're not taking into account what we really could learn from that thing—from that thing, because there is, in most outages, very seldom a single root cause. And when we talk about it in that way, we do ourselves a disservice.
So, what I would encourage you to do is—here's another QR code—if you've not read this paper, 'How Complex Systems Fail' by Dr. Richard Cook, who is a really lovely man who died a couple of years ago, I really encourage you. It's about four or five pages, it is in English, but you can translate it, there's probably a translation available, but it's a really good notion about how do complex systems really work.
And I really encourage you to take a look at it. Everybody have it? I want to make sure you have it. If not, I can show it to you later, you just contact me through any—anyway, and I'll glad to give it to you.
So, now, just for fun, I think a better—because you're probably at this point thinking, 'Well, we use the term root cause all the time. What am I going to use instead?' because I have to talk about something, what am I going to say? And I would suggest to you that instead of root cause, it's probably reasonable to talk about triggers—what was the thing that triggered the problem—and it is probably reasonable to talk about contributing factors, because almost always there are contributing factors, there's just not one root cause. And so when you talk about root cause analysis and all that sort of stuff, you're doing yourself a disservice. And it will make it harder when you do your post-incident review if you're just looking for that one thing. It's not like a murder mystery where you're looking for the one person who has the knife who happened to go like this at the wrong time to the wrong place, right? And so I just want you to think hard about root cause as a—as a word. And whatever the translation is, and someday, please come up to me afterwards and tell me how to say it in Japanese so I can tell other people in Japanese never to use it, if you would.
Okay, now let's—let me do another thing to help blow your mind. Let's talk about root cause in this way. Yesterday, the day before the outage you had today... is that a call for me? Because I'm—I'm—I'm working here. It's okay, if just tell them I'm not here and I'll get back to them later. Anyway, yesterday, the day before your outage that happened today, what was the root cause of that—of that success? What's the root cause of it running fine in your systems? Go ahead, I can wait. Tell—tell me. If you start to ask questions like that, then you'll see that it's not so easy to talk about root causes. And in fact, there's some lovely research done—this starts to get us towards talking about resilience engineering, and we'll be talking about that later in this talk as well—there's some lovely research done in which Erik Hollnagel said, for a thing he called Safety-II, that instead of paying all this attention to the 30 minutes you have the outage, why don't we pay attention to how to make things go right all the rest of the time? Why don't we spend more effort there instead of spending all the time on that—on those 30 minutes? And then there's a lovely thing by—by an MIT professor, Nancy Leveson, called Safety-III, which says that okay, that's not exactly the best way to look at it, but we can look at it in a more deep—deep way. And so I encourage you to think a little bit about, you know, what about yesterday.
Okay, next question.
People really, really, really want me to tell them what SRE should be like in their organization when they come talk to me, when—in the many years that I've talked. They really want to understand what role should we have, how should it be, how should I think about it? And they also want to know a little bit about maturity models, like okay, what does it mean when I'm a baby SRE or have baby SRE in my—in my thing, and what is it like when I have old man SRE in my organization, right? I don't know—I don't know what else to call this. And so I like—and by the way, these little notes at the bottom that say 'hat tip' are meant to give credit to the people who originally told me about these things in case you're curious what they are. I realize I should tell you.
So, Ben Treynor Sloss, who used to be at LinkedIn, had this—had this—this model that I think is—is pretty good, where he talks about that most SRE starts out in the land of firefighting. Right? And I know you know what firefighting is here. I saw some really great stuff at the EDT talk of—thing, SRE big thing. I don't know what it's called. But mostly we think about firefighting, right? How you've got—you've got things are—you're having a real problem, you have to bring the systems up, you have to get things back better. We spend—we spend a lot of our time there. And ideally, hopefully, I wish for you that you get out of that as quickly as possible. But you fix, and things work better, right? In the best world, that happens. Doesn't always happen. I know it. I'm sorry. But then often what happens is people say, 'Oh, well I fixed it. I'm never going to let that happen in production again. I'm going to stop people from running the wrong things in production. I'm going to be a gatekeeper. I'm going to tell you what you can and can't run in production.' And I'm here to tell you that this is a bad role to stay in, because nobody likes somebody who tells them no. When you get into the airport and you see the customs person, are you happy to see them? Yes, no? Not so much. Yeah, so don't be that person. You don't want to be in that role for any length of time if you can.
Instead, you might be moving on to advocate, in which your job is to go around the company or the organization and be the person trying to advocate for reliability and reliability practices. And ideally, you're able to make a change in—in that world. And this is how this might move on. And then if you get really good at it, people want to partner with you. They want to sit down with you and they want to make—they want to make their roadmaps with you, they want to build reliability to what they're building, they really want to come see you, and they're really happy to—to work with you at that level. Maybe this never happens, maybe when I say this to you, some of you are thinking that this is like a Miyazaki film, you know it doesn't really happen. I don't know. But I'm here to tell you that sometimes it can happen. And then, Ben Sloss says, 'Well, then eventually we all become engineers and everybody thinks about reliability and everything is great.' Now, I don't know if that ever happens, but the problem here with this model is that with all, you know, this notion that we might be going from thing to thing to thing, it's not a linear model. This is not the way the real world works. It is never the case that you move there. It is very likely that at some point in time, somebody's going to introduce a new system and you're back to firefighting. Right? No matter how much of a partner or an engineer you are, you're back to firefighting. And that's okay. This is the way the world runs. We're in this business, you knew this business, you knew this—you knew this when you got into this business, so I'm not telling you anything to be surprised at, that from any stage you can go back to firefighting, and it's okay. But ideally you get better at it every time, or you build the things like platform engineering to make it less likely that you're going to be firefighting, right? So it's entirely possible. And you still don't ever want to be in this in this role here. Like, you—if you can avoid gatekeeper, you should.
Okay, so the answer I think really in this case is you kind of want to ask: 'What role should SRE take in my organization today? What does it need today?' That's the answer, to my—that's the question you should be asking, not like, 'What do I want to be when I grow up?' right, which is often what people are asking when they're asking for maturity models. That's my—that's my opinion.
Okay, next question.
People want often to understand how they talk to the other people in their organization about SRE. Maybe you want budget, maybe you want to hire more, maybe you just want to stay employed, maybe you want to figure out some way to—to have SRE be a bigger thing in your stuff, and so you're trying to figure out, 'How do I tell people we should do SRE? How do I sell it internally?' And the problem is, is that we almost always do this wrong. So I want to suggest ways that people do it wrong and suggest that you don't do it those ways.
The first thing that I suggest you don't do is become an insurance salesman, where you're like, 'Here's the thing: if we just pay some money in SRE, there won't be a problem, reliability—reliability will be good.' You know? And you say to people like, 'You wouldn't want to have a downtime because downtimes are really expensive, right? So pay for SRE because they're really expensive.' It's like—it's like bad insurance, right? It's—it's unmaterialized prospective risk where you're saying basically, 'You know, nice system you have here, it would be a real problem if something happened to it, don't you think?' and you hope that they give you money. Don't do that.
Instead, it's far better to say, 'Hey, we have this downtime, it took us some work to recover from it, and then it took us even more work to build the reliability so it didn't happen again. Here's how much that cost. How about we say let's pick a number and let's use that cost as the basis of our number?' So, you're an engineer. Do an engineering job with that.
It is really, really, really easy to sell a fantasy to management. They would love it if you would simply say to them, 'Just put a dime in the jukebox and you'll get reliability, and six months from now, just put another dime in, and you'll still have reliability.' They would love to hear a story, and you know if you're in a meeting room and the bosses are nodding and they're looking happy, that you're doing it wrong, often. I mean, maybe you're better at your job, but—but I'm just telling you that if you sell this fantasy, it's not one that you're going to be able to provide because that's not the way reliability really works.
And the other thing that we also talk about is making it a tax. Like, 'Okay, make sure you spend enough money on reliability like enough money on security and everything will be fine,' right? And it's on top of the things we're building, right? On top of that, we need to have reliability. 'Oh, that's, you know, I don't know how many yen, okay, great, there's your check, you know, away you go,' you know, but that doesn't work that way.
Is there anybody in this room who knows you can't add reliability later and have it go as well as you hope? Right? Is there anybody? Please raise your hand and we'll have a lovely conversation outside in the hallway and I will tell you the world, right? It's like security, right? If you say like, 'Let's add security later,' does that—has that worked for you, anybody? I don't think so.
So, instead, you have to realize that it's a feature like any other feature in the system you're building, right? And so that's where working, you know, being part of the roadmaps is important.
The other thing that people like to do is they like to claim or they like to hope that, 'Hey, we're going to bring SRE in, they're going to help make things reliable, everything is going to be great, you won't even notice that they're there. You can barely tell that they're working, right? Because you won't have to do anything differently.' And I'm here to tell you it's like going to the gym. It's not painless, it's not invisible, and in fact, they won't—they might—they probably won't like it originally when you say, 'Hey, now we're going to do things a little differently because if we don't do it differently, we're going to just keep on having these outages.' So, when you say we're going to do something differently, they might have to be a little unhappy. So don't tell them it'll—it'll be that easy.
And then, finally, the other thing that I think that people do is they go to a conference like this, and they get so excited for SRE that they go back to their office and they talk to their boss, and they talk to the business people in terms of SLIs and SLOs. They're like, 'Oh, I love SLIs and SLOs, let's—we're going to do SLIs and SLOs.' And the thing to understand is that the business people care about business things. They don't really care about your SLI and SLO. They just care: is the cash register running? They just care, you know, among other things—I mean, that's not completely true, but you understand, right? And so you really want to go talk to them in their language, just like I'm doing a bad job of talking to you in my language. But you understand, right? And otherwise you won't—you won't understand.
So, so instead, what I want to suggest to you—because at this point you're, this is like telling you that if you want to take and make Michelangelo's David, you know the statue, all you have to do is get a big block of marble and chip away everything that's not David, right? Do you know that process? That's not what I'm telling you how to do. Let me tell you how I think you should do it. I think it's just basically, you know, give the truth based on real data in a language they speak. Right? It sounds, whatever.
Some people are taking—some people are taking pictures of this slide. It's interesting to me. I hope this is helpful. I'm more than happy to talk about any of these things much longer, you know, when we have more time to talk. And I'm also happy to give you my slides. And then you can translate for yourself. I have no problem with that, I want to help everybody.
Okay. Did everybody get that? Everybody cool with that? Yes? Any more pictures? Selfie? Do you want me to do a selfie with the slide?
Don't worry, people to my right, I'm going to come over to you next.
Okay, so now I want to talk a little bit about toil. How do we automate toil?
And when I talk about toil, I'm using the SRE definition of toil. It doesn't mean just the things we hate to do, right? I'm not using the colloquial version, again. I need to learn these—these in Japanese. But this is my way of, for fun, inserting instead AI for this. Let's say, 'How do we automate AI away?' Use AI to automate our toil. Now we get to talk a little bit about AI, but I—I want to talk about AI once I define toil, in a fundamental way. I don't want to tell you my opinion—this could be a whole other talk if you want—but my opinions about where it's going to go and how great it is and etc. I want to talk about some of the questions that will help you understand how to think about it when it shows up on your doorstep, even though it's already there.
So, just to be clear, we're talking about the same thing when we're talking about toil. This is the definition from Vivek Rau in the SRE book and in the SRE workbook, in which he says that toil are the kinds of things that we do every day that are manual, that are repetitive, that we probably can automate, that are tactical, not strategic, and the thing that I think is really important about the definition of toil is it has no enduring value, meaning if I did it today, the system isn't any more reliable tomorrow when I'm going to do it again. And then when I do it the next day, it's still not any more reliable, nothing has—nothing has improved, I just had to do this.
And then this O(N) with service growth is really good for, let me see, how many people in this room have a CS degree? Have a computer science degree? Let me just find out out of curiosity. Yes? Some. Some? A few. I also. So, what this is trying to say is, is that SREs realize that if you have something that takes one person to run the service and it gets 10 times bigger, you can't hire 10 more people. And then when it gets even bigger, you can't hire 100 more people, and you can't hire 1,000 more people—it's got to scale sublinearly. That's what that's trying to say, because there's only so many of us, right, and there's so many different possibilities. So, okay.
So, let's talk about my clarifying questions we ask about AI as an SRE.
Okay, the first question is: 'What do you mean by AI?'
Now, I know that some of these people in this room remember AI before it was just large language models, before it was LLMs, right? You remember machine learning, anybody? Anybody remember that? Back in the—back in the days when we could do AI or something like that. So, most of the time when we're talking about AI, we're talking about large language models. So I just want to be clear about that. I'm talking about LLMs now.
Okay, so then the next question I think is super important is: 'Who is using it?' And what I mean by that is, people like to—especially management likes to fantasize—that if I were to give AI to a junior person, they would go 'Poof!' and become a senior person, right? Have you seen this at your job, where the—where people where people don't really quite understand that it's who is using it that matters? Because the reason why AI can work so well, especially for senior people—and some of the senior people in this room—is you know when to argue with the magic talking box.
But somebody who's junior doesn't have the experience to know when to argue. And so it's really important to understand when—who is using it and what is really physically possible.
And then the question becomes, like an SRE: 'When does it succeed? What is it good at?' You know, 'When does it fail? And how do you tell when AI is succeeding, how do you tell when it's failing?' Right? This sounds like a very simple question, but I'm here to tell you that anybody in this room who's used it to build things knows that you didn't realize it was failing until a little farther down than you preferred. And so what do you do about it?
And what is it good at, right? It's not—it's really good at saying, 'You are the best programmer who has walked the Earth. I cannot believe that I could help build something for you.' It's really good at being sycophant, at being a sycophant. I don't know if that will translate. But it's not good at knowing things. So, if you—it doesn't have any real domain knowledge. It's fabulous at picking the right token, right? That's what it does. And that has some really nice properties about it, but it doesn't have the domain knowledge you do. It doesn't value the same things as you do. If you value security, it doesn't know to value security unless it's in the training or somebody's told you that, told it to do that. So I think this is something that gets forgotten a lot.
And then how good is the data it's been—it's been trained on, right? Because if you think, 'Oh, my junior people will become senior people because I'm going to train it on my documentation,' I don't know whether documentation in Japan is as good as the documentation in the US, but if it is, you know this is a bad idea, right? You understand that sometimes what it's trained on isn't very good or is—is a lie.
And I, one example that I—that I like very much is, I know somebody who is an educator, and this is an educator in New Mexico—it's a US state—that requires you to take classes about New Mexico history. And in this New Mexico class, this person assigns an assignment about New Mexican history and says, 'Use whatever AI you want, go ahead, totally fine.' When they do it, then the next assignment is: 'Go back and fact-check your assignment and make sure that it's—it's actually true.' And it turns out that there isn't a lot of training on New Mexican history in the major AI models, and so it makes stuff up.
And so, that's what I'm trying to say is, how good is your data?
And then the other thing I think is really important to understand is, as SREs, what do we already know about automation? Right? We've been here before, right? We know something about automation and how it succeeds and how it fails and how to tell the difference, right? We've been here. And so, I want to say, not only we've been here, there's really good research. This is another paper I really, really recommend you look at that was written a long time ago about automation and how automation can go wrong.
And when you read this, if it doesn't make you feel a little creepy because it sounds like AI to you, I would be very surprised.
Everybody got it? If not, again, I can give the slides and the stuff like, but it's a really lovely thing and it's—and it's really well done. Okay, great.
Let's do the last part here, I have another couple of minutes left.
This is my attempt to help you understand the difference between two words. Please tell me that people use the term resilient in Japanese. I don't know what the word is in Japanese, but that—and they use both reliable and resilient as two different words, yes? Do they also make them as if they're the same word? Do they use it in the same way in the same sentences? Yes? Yes? Okay. What I want to tell you is that one of the things that we've learned through a field called resilience engineering is that resilience is not the same as reliability. And in fact, what we can learn from resilience and resilience engineering is, in many ways, the future of reliability.
Oh, I would gladly take a selfie with this slide shot, because I spend all my time telling people not to use resilient when they mean—when they mean reliable.
So, and why does this matter? Well, in my experience, when people say resilient, they mean fault tolerant, redundant, preventively designed, highly available, or self-healing, right? That's what they really mean.
And the thing is, is that the difference between resilience and, say, redundant is this—this is an example from John Allspaw: if you have a spare tire in your car, that is redundant. But if your car breaks down because the battery goes out, that's—that's not going to help you. What resilient is, is knowing how to get another car, or knowing how to call a cab, or knowing how to get an Uber, or doing something else to still do what you have to do with that car. And so, resilience can mean a lot of things. It can mean—and this is from another paper—rebound, how do I get things to come back the way they were; it can mean robustness, how can things be okay when it's under load; how can we be extensible, how can we handle surprises when they happen; and how can we be okay with all kinds of surprises that they happen, it's all about adaptive capacity, right? And so, somewhat, thing I want to suggest to you is some of the future of reliability is in adaptive capacity. And the only way this works is if we use resilient when we really mean the word resilient, right? And probably your place where you work stops right there at that line. I bet your systems aren't necessarily gracefully extensible, extensible, or sustained adaptability, but they could be if you start thinking this way.
And this is the name—this is the paper that I named this talk about, called 'Resilience is a Verb' by David Woods.
And it's where I was getting that—that idea from.
Okay, my friends, I'm going to let you down. We are coming very close to the end. I'm sorry, I realize I'm going to be a couple of minutes late, but this is because I'm speaking slowly for translation.
So, now we come to the end. My goal is to help you think a little bit about the questions that you're asking, how to make them better, how to think about them. Reliability, to my mind, is about the questions we ask. These are some of the questions we asked today. I think you probably have questions. If some of your questions included 'How could I show David around Japan for the first time on Sunday or Monday?', that's good, that's a good question. Or if you're thinking, 'Gosh, it would be really great if I could get this guy to work with me, or to come to Japan to talk to my company,' well, good news! Good news for somebody.
I'm hoping that you'll come ask me questions at the Ask the Speaker, which is where I'll be after this, and any other time. So with this, thank you.
**[41:35]**
**司会:** デビッド・エドウェル・ブラント・エデルマンさん、ありがとうございました。このあとはAsk the Speakerを行います。セッションについての質問や感想はAsk the Speakerでお願いします。こちらの会場は会場転換を行うため、一度ご退室をお願いいたします。