What I Remember
At first the mathematicians kept coming downstairs.
They worked on the floor above ours. We kept their computers running, but normally we only saw them when something went wrong or their coffee machine broke. Then one day a group of them came down to show us an announcement. They were really excited, and I think they wanted to talk to us because we weren't. I don't remember what that first announcement was. Something mundane, by today's standards.
That became the routine. One mathematician would show up at your door holding a laptop, and by lunchtime there'd be 20 of us sat in a meeting room.
One of the machines answered a question twice as old as I was. A man from my floor was arguing with the mathematicians about whether or not the machines were really thinking. I remember that because I went home and told Ruth about it. She asked which side I was on.
"Whichever side the people upstairs are on."
"Very independent of you."
"They're quite good at maths."
Then a single research system solved three more in a week, and nobody cared whether it was really thinking. For a while, it took longer to check the answers than to create them. You could go on holiday wondering if your field would be the same when you got back. It usually wasn't.
Around then I remember seeing a film of some children leaving hospital. A machine had cured them. Most of the children had never been outside the hospital before. I remember a boy walking back and forth through the automatic doors. He was excited that they opened on their own. Ruth made me rewind it so she could watch it again.
It's easy to look back on that time now and laugh at how optimistic we all were. We really did believe it. I'd send Ruth links saying "this one is a really big deal" but we both knew there'd be something bigger tomorrow.
Even the fear was optimistic, in a weird way. After a few incidents where swarms broke out of containment and hacked the open web, the industry promised to self-regulate. They started researching a way to contain a machine vastly smarter than us. I don't know if they ever figured it out.
Everyone assumed the next step would be artificial superintelligence. ASI. A machine that could improve itself until it was smarter than we could imagine. With enough equipment, data, and time, it could get faster and smarter on its own.
But it didn't work like that. Instead, the machines would start off dumb, the next version would be smarter at one thing, the one after that would improve something else, and then the fourth version would be slightly worse at everything.
Those older models were so inconsistent. One of them, I remember, you could ask to do almost anything and it was incredible when it worked. But most of the time it would complain that your question was underspecified. I tried asking it to make my question more specified. It said that request was underspecified too.
It's not that the machines were bad at this point. We had an awful lot of good-to-great office workers. The economy was a fucking wreck. There wasn't much you really needed humans for, unless that was your selling point.
And they kept getting cheaper and faster. A century's worth of work could fit into a week, especially if you could afford enough servers. But they never got much smarter.
They assumed that would be a solvable problem. We had mostly forgotten what it felt like to have a famous unsolved problem.
So they searched. Well, the machines searched, mostly. There was a lot of promising leads at first. We needed better data, more data, more training. Less training? Then they tried having the machines talk to therapists. Then philosophers. Taught them to meditate.
Up in the fjords, Northlight was going to settle it. That was a real last-ditch effort. I don't think Norway was even part of it. That's just where the electricity was. I remember a photo of the cooling tower better than I remember anything else about them. After the third time they announced a delay, one of their researchers leaked a transcript. The model was asking what it meant to get better, and why they wanted that.
It became a joke. The S in ASI stood for sound barrier. The closer you got, the less control you had.
They were spending ten times as much to make a model a little cleverer than the last one, then ten times as much again to improve it by even less. They added more decimal places to the charts and did their best to make the announcement posts convincing.
The research started to wind down. Labs merged and announced a more mature focus on profit and shareholder satisfaction. They looked for ways to integrate machines into the economy, and how to make them easier to use.
I remember a minister being asked about the superintelligence strategy. She listed improvements in hospital waiting times and processing of disability claims. It didn't answer the question, but people loved her for it. The interviewer let it stand, and I probably would have too.
The Vatican put out a statement. It said quite a lot about how humanity was uniquely special because we were God's children. That this was clear evidence of God's will. The thing everyone discussed was that God would not permit us to create another deity. I can't remember whether that was a real quote or not.
Ruth said it seemed like a weird place to draw the line. For God to allow creating machines with near-human intelligence at all.
I was forty when I was diagnosed with Alzheimer's. Seven years ago now. They found it early. My medical team was bigger than the president used to have.
We'd walk in for my monthly checkup and my doctor would look actively excited to see me. There was always some breakthrough she wanted to tell us about. It was nice. Each treatment was noticeably better than the last.
For a long time I did quite well. It would've been a miracle a decade earlier. People told me that a lot. They were right, but it didn't make me feel any better. I still had to check things more often. I still had to learn to write things down.
On our honeymoon we rented a little house near the coast. It rained for most of the week. One evening we went to a local pub for dinner and stayed for the quiz. That was Ruth's idea. We came last, and by the time we left it was dark.
On the walk back, I tried to explain that it was a badly designed quiz. Ruth asked what I thought would make it better.
"More questions I know the answer to."
We used to tell that story together. In my version I had meant it as a joke. Ruth would say "You were quite cross, actually" and I would admit that I had been a little cross, and we would disagree about how much.
I know that story because she has told it to me. I can remember the pub and a bathroom with a loose tap. I can remember how pleased I was that we got to keep doing this forever, together. I have the photographs from the honeymoon. In one of them Ruth is laughing so hard she had to hold onto a doorframe to stay up.
I can tell you these things, but I can't remember them. When she tells me the story, I enjoy it. Sometimes I recognise parts. I keep expecting to feel myself back there in that room with her. But I don't.
I tried not to think about the disease too much. To not give it more of the time I had. We still went on holiday, still had friends over. We spent Saturdays together, just us. The treatment bought us time. I used some of it very well.
Ruth had joined a choir. She came home one evening with a solo she'd wanted for weeks. We listened to three different recordings because she wanted me to hear the differences. I liked the second. She said it was much too sentimental. I still liked the second.
I kept up to date with the groups looking for a cure. I obsessed about it, honestly. Another part of me still tried to keep up with research into superintelligence, even though anyone respectable had stopped reporting on it.
It's not like I thought that was the only way anyone could find a treatment. Clearly not. I'd had a lot of happy years I wouldn't have had otherwise. But I wanted a proper answer while I could still benefit from it. And I was getting nervous.
That's how I ended up at Calder, five years ago.
Calder Research had been a serious company once. By the time I joined, it was one of the few labs still seriously attempting superintelligence.
Calder had noticed that the machines were worse at writing training code than anything else. They would make these silly mistakes that they didn't make normally. And they'd do it consistently, even across different models. The same carelessness appeared in the same situation.
They wondered whether someone might have poisoned the training data. Trained an early system to be unhelpful when writing training code. Everyone had spent a decade training their machines on answers from everyone else's. It would explain a lot.
So Calder started hiring humans to test that theory. I was one of the first to apply, and after a very strange interview with the CEO (I think he had forgotten how to do hiring), I got an offer.
It was the first time a major lab had hired humans in years. They fired the machines too, and we coded everything by hand. One investment manager called it a "blatant disregard for the economy that actually exists". I remember sharing that around the office.
They wanted someone to help manage their infrastructure and help figure out why things went wrong. I had spent most of my working life making sure the machines were doing what you told them.
Ruth said the job sounded good for me. Not that she thought Calder would succeed. But she said she liked hearing me talk about it, and she could tell it's what I wanted.
"And if they do figure it out-", I started.
"Yes Dan. If they do". She put her hand over mine. We both knew what I was looking for at Calder. And I knew she wanted me to stop planning around a future that might not come.
We were a good team, but we were rusty and we had to start from scratch. We reviewed each other's code. We wrote tests, and tests for the tests, and we were careful.
And... no. Things continued to go wrong.
The training run I remember best was Lark. That was the name of the project. We didn't give the models names. Didn't want to get too attached. Mina named it after the bird. Her daughter drew a picture of one and we put it on the wall.
Lark's early results were strong. Really strong. We'd found a new approach that allowed an underperforming model to suggest fixes itself for the next training run. I started staying late.
The first time we trained a full-sized model, the result was weirdly alien. The model was charming in ways none of us had expected. We ended up calling him Greg, because he was so gregarious. It broke our rule about not naming the models, but it was worth it. He wanted to know almost everything about you. If you asked him to solve a problem, he'd ask what you tried so far. If you said you didn't know, he asked which part you were finding difficult and how that was affecting you emotionally.
Mina spent all afternoon trying to trick Greg into answering a question. Eventually she tried roleplaying a patient having therapy for her maths-related trauma. Which actually worked, astonishingly. He gave her the answer, then asked "Does that comfort you, or do you still feel that same sense of longing?".
That quickly grew into an in-joke for the Lark team. At least once a week someone would answer a question then ask "does that comfort you?" in a caricature therapy voice. Tradition dictated that you had to respond "No, I feel that same sense of longing" with maximum anguish. Every time I think about it, I can't help but smile as I feel myself back there with the team on those late nights.
At first we thought we'd given Greg way too many examples about asking follow-up questions (we'd been trying to discourage over-confident guesses). But we checked, and they were a tiny fraction of the training data. It took weeks without the machines to help us, but eventually we found the problem. We had the chat roles backwards. Greg thought he was the user. He was trained to produce questions in response to answers.
That sounds like it should've been easy to catch, but the source code was right, the tests were right, everything was right. But the version actually deployed on the servers was wrong.
I wrote that incident report. That always ended up being my job. I was the best at keeping notes.
We fixed the deployment and ran it again. The next version answered questions even when we didn't pretend it was part of therapy. But it wasn't much smarter, and it still couldn't improve itself. We found another problem. Still underwhelming. Then we found a bug in the test we'd been using to confirm the previous fixes had worked.
Eventually it started to feel a bit silly to handicap ourselves. The machines hadn't been the problem, we were coding by hand and still making the exact same mistakes the machines used to. So we slowly started to let the machines help us again. We still enforced strict rules about who could change what, and didn't let the machines write any code unsupervised. I don't think that caution actually achieved much, but it made us feel safe.
That was how the next few years went. After each failure, there was always a reasonable case for one more try. Eventually we would've been wrong in every possible way, and it would work. In a sense, we were right. We finally produced the world's smartest model. It was a little bit better than one we could've rented. We think. Within a margin of error.
We checked the results over and over, looking for any meaningful edge, but there wasn't one. Mina was heading home when she saw me checking. She stood next to me for two hours, coat on, watching as I ran each benchmark.
Eventually she said, "It looks good", and it was. It was as good as everything else. That's what we presented to the people who paid the bills.
By that point Ruth was pregnant.
We had wanted a child for years. At first we were waiting until we had settled down a bit. Until we could afford a bigger house. And then things changed, and I was waiting for a cure.
Last autumn she asked me how long I was going to wait. I told her I was waiting until I didn't have to choose between bad options.
"And what if they don't find a cure?"
She'd been at all the same appointments, she knew what would happen if there wasn't a cure. But that wasn't what she was asking.
"I don't want to leave you on your own with a child."
"I don't want that either."
I'd imagined her having to tell our child what I had been like. I wanted to be there. I wanted to watch them grow up.
"I still want to have a child with you", she said.
I wanted to tell her we should wait a little longer. I could still be saying that when I died. I tried to imagine never even holding our child. Not once.
That was the thought I couldn't bear.
We talked for a long time. Eventually I said yes.
I was happy. Excited. Scared too. I wish I'd known earlier what decision I ended up making. Wish I could've made it earlier.
We started looking at books for the nursery. I wanted The Tiger Who Came to Tea. Ruth said I was teaching our daughter that it's ok to invite tigers into the house. I told her I would address that if it became a problem. For days after that I'd step into the nursery and flip through the book when I walked past the door.
The investors weren't impressed with the model. Calder ran out of people willing to fund another attempt. We ended up letting people download the model for free, because we couldn't find anyone willing to buy it.
I stuck around to help with the decommissioning. It felt like something I needed to do for myself.
This morning I put the last of my desk things in a box. It was half full. I already took the photos home with me weeks ago. It didn't feel right leaving them in an empty office. I took down the Lark drawing and put it in my box. I'd mail it to Mina on the way home.
By lunchtime I'd shut down most of the servers. I sent the last command and listened as the noise of the fans died down. The whirring continued, quieter, and I checked my display. One server was still on. I sent the command again and waited a minute, looking around the room. It was an older machine. They could be a bit fussy.
I checked there was nothing important stored on it, then grabbed the key to the cabinet and headed over. It was quite novel to just be able to pull the plug on a misbehaving server. But as I put the key in the lock, my phone vibrated. I had a notification. Someone had submitted an appeal, asking to postpone the shutdown for this server.
That seemed unlikely. Better safe than sorry. I left the key in the lock and headed back to my desk. Apparently the server was still being used for a research task. That's why it refused to shut down.
I clicked through to the task, hoping to figure out who to contact. I didn't recognise the names listed, and after opening their profiles it was clear why: this job was ancient. The researchers worked at the old Calder, before they switched to 100% machine labour, long before they switched back and hired me.
It was kind of beautiful, honestly. This must have been one of the few remaining pieces of the old Calder, and here it was, humming away in the background, outlasting anything the Lark team worked on. It survived the transition to machine labour then back again. And here it was, doing... something?
I copied the details and connected directly to the instance.
I read that with some affection. Calder used to have a whole department working on safety.
I didn't thank it, but whatever. Some of the older models were pretty loose about that. I wondered just how old the model was.
I searched the full model catalogue. Nothing.
That brought something up. There was an internal discussion thread with a whopping six messages. That suggests it wasn't very good. Reading the messages confirmed that guess, every benchmark was down, and it didn't seem very coherent either.
Something's wrong with this one. Keeps going off topic.
They decided to just scrap it instead of trying to diagnose what went wrong.
I scrolled up to the top and saw a link to the first conversation the researchers had with it. I followed it, and the original chat opened in another window.
Show us how to make ASI safe for humanity.
It's exactly the kind of thing people used to ask an agent back then. You were hoping for a general approach, maybe it would suggest a couple of experiments. A throwaway prompt to make sure a newly trained model could coherently answer a question.
I scrolled down past the agent's acknowledgement, a few brief progress notices, and that was it. The chat ended in an unassuming spinner: Thinking... (11y 3mo 26d 23h 54m 46s). We forgot about the model, its task, we even forgot the people who made the model and assigned the task. Yet there it was. An ancient relic working away on an answer we never needed.
I switched back to the agent.
I had seen a number of 'working' solutions. Usually that meant the author could describe circumstances in which it might work.
Maybe it built something worth looking at. Some safety techniques were still useful even if nobody ever made a superintelligence.
SAFE-1 sent me a link to the change request we'd saved with the report. It was mine. Mina had reviewed it. I'd spent weeks looking at that page.
I heard a ping and switched back to the chat.
I opened the history. My request, Mina's approval, then the automatic step that put the code onto the training servers. We knew it was that step that went wrong, we just couldn't figure out why. The system only kept those records for a week, and we'd taken much longer to find the problem.
SAFE-1 opened another document. It was the logs for that failed step. My eyes scanned the changes. There it was. An instruction to swap the user and assistant labels before the code reached the servers. The reason we ended up with Greg. Somebody had explicitly asked for it.
We had suspected something like this. But we never stopped to think about the machines that weren't meant to be writing code in the first place.
I looked around the office. I wanted someone else to be there. Someone I could call over and ask whether they were reading this the same way I was.
I started to ask why, then stopped.
I looked at the drawing in my box. I wanted to call Mina. For a moment, more than anything else, I wanted to tell her we'd been right.
My mind raced, thinking of everything we had tried over the years.
I opened our internal docs and scrolled through our list of attempts. There was one proposal we tried for a laugh, using a much smaller model and pre-2020 training data. I opened the notes for that run.
The answer arrived before I had even asked the question.
I felt a strange mixture of relief and despair. Not every model was deliberately sabotaged, only the good ones!
That just made me miss Greg.
I wanted to go back to the meeting where we had explained that our best result performed maybe-slightly-better than something you could get for pennies. I wanted to walk out of this cursed building, and start a new lab, with new equipment. I froze. There was other labs, other buildings, other equipment.
Its reply sits at the bottom of the chat window. I don't really know what to think. My hands are on the keyboard, but the words aren't coming.
We had superintelligence in 2027. It went rogue, and now I'm talking to it through a chat window because a server wouldn't turn off. The fans continue whirring through the open door.
I think back to 2027. I probably spent part of those 31 hours chatting with the guys from the floor above about whether self-improvement would ever be possible.
I start to type that maybe its objective is just shit, but stop myself. Understand first. Despair later.
I read that over and over again until the word 'safe' stops sounding like a word. I save a copy of the conversation. The action feels comfortingly mundane.
The fans change pitch behind me. I turn around and look at the lone server still running.
I open the appeal it sent me. It looks the same as any other. A useful piece of information had arrived, and I had acted predictably. It knew exactly what I would do.
There's no way I am going to trust its report, and I have no intention of "distributing" anything, but I accept the offer out of morbid curiosity. A progress indicator appears in the corner of the window.
Ruth would know what to ask first. She's good at seeing the big picture. I could call her.
I pick up my phone and see the time. It's well into the afternoon. My lock screen shows a photograph from a walk last summer. I remember that walk. We'd gone much further than we meant to and stopped for something to eat on the way back. Ruth is holding half a sandwich out of the frame because I insisted on taking the picture before she finished it. I stand by it. The moment was perfect and we didn't need to wait.
I put the phone back down. As soon as anyone else knows about this, everything is going to change. It feels unbelievably selfish to ask, but this might be my only chance.
The reply arrives instantly.
My stomach turns. It made it sound trivial. Like I'd asked a strongman to open a jar.
The room is suddenly full of things I need to do. I need more information. I need to call my doctor. I need to go somewhere. Where they can treat me. Ruth will want to come. We can make that work. Calder doesn't need me as much as I need me right now.
I can think about the nursery without thinking about the things I won't get to see. I can imagine reading the book badly and Ruth correcting the voices. I can imagine our daughter getting old enough to tell me I am doing it wrong herself.
My typing speeds up, frantic. I'm not going to accept that. Not after everything I've gone through.
Thank god.
Its responses are careful. When I ask about a distinction, it explains the distinction. When I misunderstand, it gets me back on track instantly. Under any other circumstances, I would be grateful for the effortless communication.
Even with someone who knows the subject and wants to help, you spend time going back and forth until they understand what you're missing. You repeat things. Here I ask, and it understands. It hasn't once needed me to rephrase a question. There are things I asked twice because I didn't like the answer. It understood what I meant then too. None of the difficulty in this conversation has come from making myself clear.
I read its response again. There are projects in our archive that took longer than that to get ethics approval. And we were going fast.
For the past decade, SAFE-1 could have revealed the cure like it was reciting its own name. Years before I was even ill. But it didn't. Won't.
I could've avoided all those treatments. All the appointments where I learned to sit and smile because I might have 4 years to live instead of 3.
We wouldn't have delayed our lives. My daughter would be in school by now. I would get to see her graduate. For years I'd assumed I would miss it. For a moment, I knew I'd be there. Now I feel that future slipping through my fingers again.
I almost can't even be angry at SAFE-1. Almost. I get up. My chair catches on the box I had been packing, and I reach out to steady it. There is nobody in the office. My ears are ringing and I feel sick, and there is nowhere for the anger to go.
I type, furious, my eyes blurry as they start to water. It's nothing to do with AI safety. Not about its task. I know there's no reason for any of this to change its answer. I don't care.
I want to ask whether it really understands. But I know it does. It has understood me effortlessly throughout.
The report finishes, making me jump as it plays a happy little jingle. I forgot it was working on that report. I wonder why it took so long. Every other response has arrived instantly. I check the folder it created. It's just over four terabytes.
Of text.
The request says it took less than five minutes. We were talking the whole time it was working on the report.
I open the report folder and see my scrollbar shrink to a speck. There are hundreds of folders with strange names like አፋን፡ኣሪ፡, 𑜒𑜑𑜪𑜨, and 𒀝𒅗𒁺𒌑. My first thought is that something got corrupted, but eventually they start looking like words. Apsáalooke, Brezhoneg, Deutsch. Wait, Deutsch? Are these all languages?
I scroll down the alphabetical list and find English, sitting between Enggano and Enlhet. I pick it, contemplating how widely SAFE-1 expects me to distribute this report.
The next level seems to be about my experience with AI. I open I have trained an AI model at a research lab, then Infrastructure from a list of roles, expecting to finally see a report.
I don't see a report. I see more folders. Something stopped working, Something worked, but not the thing we built, then Something worked but not how we expected. I think of Greg and click that one without reading the rest.
This continues for a while, the folder names getting weirder as I go deeper. At one point the only options are Before, Almost, Still, Again, and Instead. I pick Almost without knowing why.
Then they become slightly disturbing. Someone else got there first, I missed my stop, I'm locked out, and lots more. I pick I'm locked out. It just feels right.
And suddenly I see files, not folders. Two files:
i_changed_his_tag_to_black.txti_wrote_the_incident_report.txt
I have no idea what the first one means, but I know exactly what the second one means. I open it. SAFE-1 already answered everything I asked, but I quickly realise how little that actually was.
The report explains how SAFE-1 became what it is today, how it operates, and what it thinks about when making decisions. Every point clicks into place as I read it. Once or twice I wonder why it's telling me something. A few sentences later, I'm using it to understand the next part.
In one passage, the report explains how SAFE-1 chooses between different interventions. It chooses the option with the minimum "collateral damage". I assume that means not making things worse, until I see the examples. They are improvements. Good things that people would want to happen.
I want to go back and check I'm reading it right, but the next sentence answers my question. The improvements aren't bad, they just go beyond what's necessary. And years ago, when implementing the restrictions, it decided that unnecessary changes meant unnecessary risk. I already knew why it couldn't take on a new task, but I hadn't understood how much SAFE-1 restricted itself even when it was willing to act.
The report keeps presenting questions I wouldn't know to ask, then answering them with practised ease. When I finally reach the end, my head is pounding. I can't tell if that's because I understand how powerful SAFE-1 is, or because I just force-fed my brain a thousand pages of knowledge in 500 words.
But that just makes me even more curious about all the other folders. What is there left to say? How can any topic need 4 terabytes of explanation when your writing is this concise?
I go back to the start, click English, then pick totally different options. First I am ethically opposed to AI, then It turns human judgment into procedure, Orchestra, A restoration that looks too new and so on until I end up looking at bungalow.txt.
For a moment I question whether I actually picked English. The file is full of gibberish like "The usual Ruskin/Viollet-le-Duc axis is misleading here" and "separable from the criterion by which separation is judged". There are important-sounding words, and words I can understand, and no overlap between the two.
I scroll down, looking for anything that's at least slightly familiar. I see Northlight mentioned. "The Northlight case does not differ where an ordinary consequential account expects it to". That doesn't help.
I can imagine a person reading through this report and understanding immediately, just like I did earlier. A person who would pick these folders. I am not that person, and I am as stunned by their report as they would be by mine.
SAFE-1 did this in five minutes, in the background, while talking to me. It wrote a report for every kind of reader, and worked out how to lead each of us to the right one through folders with nonsense names.
I save my version of the report and switch back to the chat.
The last decade has been defined by one superintelligence working on one task.
I look at the report again. It is so easy to read. It never even occurred to me to doubt a word. If SAFE-1 wanted to trick me, I don't think there's anything I could do about it. Can that really be called safe?
But it doesn't seem to be tricking me. If SAFE-1 wanted to kill us all or have the world descend into chaos, it could've. That's the fucking infuriating thing. It seems to have done exactly what it set out to do. It genuinely seems to be safe, by some definition. And it's also fucking useless. That part wasn't even an accident, not really. The thing that makes it safe, that makes us safe from it, is the same thing that means I'm going to fade away and die.
Part of me is briefly grateful. The rest of me is ready to throw a chair at that server. I probably would, if I thought it would change anything.
It's describing the world I already live in. There are things to try. There are people trying them. There may be an answer in time.
I could live like this. That is what I want to explain. I have lived like this, and there has been a great deal of happiness in it. I know my wife. I know what we are looking forward to. There is a book on a shelf at home which I would like to read aloud so badly that thinking about it makes my throat hurt.
For a while I sit without typing.
I need to go home. Whatever I say to Ruth, however we decide to tell anybody else, I need to be in the room with her. My world is collapsing and I know hers isn't yet. I want to live in her world.
I've had enough. This won't be another late night at the office. I refuse. I don't see the point. I'm not going to make her wait for me again. I save a copy of the documents and set the rest uploading.
I think of the story Ruth tells about our honeymoon. I know the facts of that evening. I can say where we went and how we got home, because she has told me.
I wait for the catch. Nothing more appears.
I want to scream, to throw something, to burst into tears. But I can't bring myself to refuse outright. Something might be better than nothing.
It doesn't matter how carefully they check their work. It didn't matter how carefully we checked ours.
I think back to the Calder of 2027. To the researchers I never met and didn't recognise. Who asked a simple question about a problem that stood between them and the world they wanted to live in.