Show Notes
What happens when millions of people start asking artificial intelligence, rather than their pastors or Bibles, to define the gospel?
Author and program director of the Keller Center for Cultural Apologetics, Michael Graham, joins host Case Thorp on the Nuance podcast to unpack the groundbreaking “AI Christian Benchmark” report. They explore how different AI platforms respond to historic Christian questions and the profound theological risks as humanity shifts to AI-generated answers.
Whether you are a pastor guiding a congregation or a professional navigating a rapidly changing workplace, this conversation will equip you to use AI with wisdom, intentionality, and a firm grip on truth.
📚 Episode Resources:
The Great De-Churching by Jim Davis & Michael Graham: https://www.amazon.com/dp/0310147433/
The AI Christian Benchmark: https://www.thegospelcoalition.org/ai-christian-benchmark/
The Gospel Coalition Website: https://www.thegospelcoalition.org/
The Keller Center for Cultural Apologetics: https://www.thegospelcoalition.org/thekellercenter/
Nuance is a podcast of The Collaborative where we wrestle together about living our Christian faith in the public square. Nuance invites Christians to pursue the cultural and economic renewal by living out faith through work every facet of public life, including work, political engagement, the arts, philanthropy, and more.
Each episode, Dr. Case Thorp hosts conversations with Christian thinkers and leaders at the forefront of some of today’s most pressing issues around living a public faith.
Episode Transcript
Case Thorp
Artificial intelligence is no longer a niche theology. It’s rapidly becoming a primary way people search for answers about life, answers about meaning, and even God. By 2028, it is predicted that as many people as search Google now will be searching AI, and it will be the new front door for even theological inquiry.
Well, our guest today is a good buddy, Michael Graham. First and foremost, he’s a friend that I have enjoyed over the years. He currently serves as program director of the Keller Center, the Tim Keller Center of Cultural Apologetics at the Gospel Coalition. In that role, he has a chance to shape Christian engagement with some of the pressing issues of our time. He previously served in pastoral ministry, and has written and spoken extensively on cultural apologetics, institutional trust, and the generational shifts within evangelicalism. He is co-author of The Great De-Churching with Jim Davis and working on a new book with sociologist Ryan Burge that will be coming out soon, and we’ll have him back on the show. Michael also is one of the leaders behind a new study commissioned by the Gospel Coalition called the AI Christian Benchmark. The AI Christian Benchmark, we’ll have a link to this in our notes, is a groundbreaking report evaluating how the different major AI platforms answer fundamental Christian theological questions. When I heard of this idea, I thought, my goodness, what a neat experiment, but the results, let’s see where they land. So Michael, thanks for being here.
Michael Graham
Thanks for having me, Case.
Case Thorp
Yeah, you’re just up the road and your children go to school on the same campus as mine where I work and we don’t see each other.
Michael Graham
That’s right, we don’t see each other enough.
Case Thorp
We don’t. We’ll have to grab lunch for sure. Well, friends, let me welcome you to Nuance where we seek to be faithful in the public square. I’m Case Thorp, and so glad to have you here. Reminder, please like, share, leave a comment wherever you may capture this podcast. It really helps us to reach further and get in front of other audiences. Well now, the A.I. Christian benchmark study. It evaluated seven leading AI models and their particular response to historic Christian questions such as, who is Jesus? What is the gospel? Was Jesus raised from the dead? Is the Bible reliable? And then scholars graded these responses using standards rooted in historic orthodoxy, including the Nicene Creed. And this wasn’t a gotcha approach. It was just essentially a technical audit to see where the theology lands as these early AI models attempt to represent our faith. And so that’s what I’m really grateful Michael’s been a part of and going to share with us. So tell me, Michael, I mean, what prompted, and to be clear, I mean, it is the Keller Center that called for this report. And what prompted, and where did this idea come from?
Michael Graham
So, about a year ago, I was having a conversation with my mom. And my mom’s in her 70s. And she’s of what you’d say, average intelligence. One of the things…
Case Thorp
Careful, you mean as compared to artificial intelligence?
Michael Graham
Well, yes for peers.
Case Thorp
There you go. Not her peers.
Wow, okay. I know my mom will be watching this and I think she’s brilliant and wise and beautiful and okay, go ahead.
Michael Graham
I love you, mom. So I was asking her just a little bit about her use of technology. And this is early 2025. And she was telling me about her use of technology. And it became clear that she uses Google, Google AI Overviews, and Gemini all on a daily basis. But what was interesting is she didn’t know the difference between the three. And I don’t think that that’s shameful or really even all that abnormal, especially for early 2025.
And it dawned on me, so I asked her a couple more questions about what her usage of the platform, those different platforms looked like. And she basically treated all of them the same way. She treated all of them as if they were just Google. And so the interesting and important thing to note there is, you know, if you’re, let’s say you’re middle aged and you’re in the knowledge economy. Well, you know, there’s the entire field of prompt engineering, which is where you give AI models a lot more context than just like a basic search of like, hey, what is the Gospel or did Jesus raise from the dead? But so, you know, let’s say I was writing something that was more thorough, maybe a sermon or something. And I wanted to know, hey, you know, what was John Owen’s perspective?
Case Thorp
Give us an example of one of those.
Michael Graham
You know, on a particular chapter or, you know, or passage of scripture. Well, yeah. So I would give a lot more context of like, I’m doing this kind of work, here I’m preparing this type of document. I’m doing research on this particular passage. And I’m very curious about what John Owen and other similar Puritans thought about this particular…
Case Thorp
Puritan theologian.
Michael Graham
Passage of scripture, please make sure that you give me responses that are consistent with these particular creeds and confessions of faith. Nobody, that’s not how keyword-based search works.
Case Thorp
Like normal Google.
Michael Graham
Yeah, like normal Google. So normal Google works based on basically keyword ranking. How semantic search works is totally different. Semantic search is the, you know, that would be the technical term for like AI or large language model based search. How some like, like ChatGPT, Gemini, Claude, these kinds of platforms. So in semantic search, what it’s looking for is linguistic clusters. Okay. And clusters that have statistically high probability of kind of word clouds. So semantic search is a lot more powerful because…
Case Thorp
Give me an example of a logistic cluster.
Michael Graham
Yeah, a linguistic cluster would be, if I searched for, let’s say I wanted to search for penal substitutionary atonement, right?
Case Thorp
Okay, and these are great theological ideas. For the lay person, maybe they just wanted to know what Mother Teresa thought about something versus Billy Graham.
Michael Graham
Yeah, so like penal substitutionary atonement is the idea that, you know, Jesus died and paid for your sins. So that would be something that somebody who’s like a pastor, you know, would search if they were preaching on a text that dealt with Jesus’ life, death, and resurrection. So that would be a, that’s a very specific keyword. Now, what when in semantic search, if you search for that, you would really only come up with articles and stuff that would, you know, be about penal substitutionary atonement. But if you searched for something like Jesus’ death and resurrection in semantic search, you would come up with stuff that dealt with penal substitutionary atonement and dealt with other types of redemptive themes like Christus Victor and these different kinds of things. In other words, semantics, you know, keyword search is very linear and it’s wooden and it gets you exactly what the keyword is that you were looking for. So for example, let’s say you’re on a website like the GospelCoalition.org, TGC.org. And right now our search function is not good because it’s based on keyword search. You put in a keyword search on our site and you’re going to have a really hard time finding, you know, kind of exactly what you’re looking for because it’s wooden. It’s trying to draw a straight line between an exact keyword and articles that have that exact word in it.
Case Thorp
Yeah. Well, let me just give an example. So my name being Case Thorp, if I Google my name, I get all sorts of websites from legal documents that end with the word Case period and then Thorp. That some other poor guy and all his legal documents, the legal Case period Thorp contends that da da da da da, right? It’s going for those direct word to word.
Michael Graham
Yes, exactly. Yeah. It’s just, yeah, it’s looking for a one-to-one relationship. Right. But what semantic search does is it’s, it’s kind of like, you know, kind of like keyword searches and maybe we can even edit some of this stuff out. One way to think about it, like keyword search is like, let’s say you’re on stub hub and you’re looking for, you know, and you’re looking for a very particular seat that you’re going to purchase, you know, at a ball game, a show, you know, so on and so forth. That’s kind of keyword search, you know, whatever section 101 row 13 seat A, you know, that’s what keyword search does. But what semantic search does is it gets you into the ballpark. Like, show me all the seats that are in, you know, in this part of the stadium or, you know, in mezzanine three or, you know, whatever.
Case Thorp
And you don’t mean technically stub hub. You’re talking philosophically that when you use semantic clusters, it’s taking you into the entire stadium of ideas for a particular subject, right, that are around those keywords, okay?
Michael Graham
Yes. So there’s just a lot more power that’s there, you know, behind all of that, because these platforms have ingested basically every, the entire internet, and all non-copyrighted material. And so, how basically large language models work, is it’s really just as simple as two things. It’s words plus statistics, words plus statistics.
And so what the large language models do is they notice in the trillions and trillions of words that they’ve digested, that there are patterns of words that appear in close proximity to each other. And the proximity that they have to each other develops, you know, linguistic patterns or semantic patterns. And so there’s no actual reasoning that’s taking place in a large language model.
It’s just statistics. It feels like reasoning, but that isn’t what’s occurring. It’s just that, you know, there’s this nerdy thing in computer science called Moore’s Law. And Moore’s Law states that computational power doubles every two years. And so at some point, if you’ve ever played the doubling game, you know, after you double, you know, 20, 30, 40 times,
Case Thorp
It feels like reasoning.
Michael Graham
You know, what you’re dealing with is just astronomical amounts of compute. And we finally hit a tipping point with computational power and energy and water and silicone and how small that they can print, you know, stuff on silicone wafers that the kind of compute needed to make it look like reasoning has finally hit a kind of tipping point.
Case Thorp
Okay, so this is really interesting on the AI front even. And I have subscriptions to all three of those, but honestly, out of your recommendation, because I don’t know if you remember at lunch one time, you were like, but Claude is so much better on this and this and this than Chat. And I’m going, is it, really? Why? So I use all three. But I do tend to Chat more, and I don’t know why.
Okay, but bring us all now to the church and to theology and ministry. What theological risks are there if Christians ignore this shift?
Michael Graham
So there’s a couple different overlapping issues and there’s a couple layers and I’ll try to kind of address each of those one at a time. The biggest shift that I think it’s important for us to recognize is humanity as a species is moving from acquiring knowledge from primary sources to acquiring knowledge from secondary sources. Let me explain what I mean by that. So primary sources would be like, hey, I’m reading a book, I’m reading an article, I’m watching video content, I’m listening to audio content, okay? And I’m doing that directly from whoever authored that content. In the generative AI era, which would include large language models on words…
Case Thorp
And to be clear, large language models are things like ChatGPT, Claude, Gemini. Okay, go ahead.
Michael Graham
Yes. And then you have image-based generative AI and video based generative AI and then audio based generative AI.
Case Thorp
Let me just throw in, as we probably all have thus far, I played around with the image AI. And I took a picture of my wonderful sister, whom I love deeply, and had tattoos put all over her face. She looked like she had just come out of a cartel in northern Mexico. And she did not appreciate this.
Michael Graham
Hahaha! What could go wrong, guys?
Case Thorp
I mean, right? Who doesn’t want to see themselves tatted up like a cartel member? Anyway, go ahead.
Michael Graham
Yeah, so going from each of us reading authors directly, in the moving from when you use traditional search, you’re still reading primary authors with Google. You search some kind of phrase and up comes 10 blue links. Maybe you click on three of them and you’re still reading primary authors.
Case Thorp
But help me, because in my mind I think, okay, if I search something on a concept of the theologian Augustine, well, there’s his primary original texts, the Confessions, but then there’s thousands of papers. Wouldn’t those papers by a seminary student be a secondary source?
Michael Graham
In a way, yes. Yeah. But the difference is in that case, when you have people who are writing scholarship on, you know, well-known figures in church history, those people didn’t read the entire sum of all text that has ever been created. So take, for example, a question like who is Jesus?
So in a situation that’s analogous to your Augustine’s confessions and somebody who’s written a paper on it. When you ask a large language model the question, who is Jesus? It’s ingested all sorts of sources, primary sources that would be in the Christian tradition, the Islamic tradition, the Jewish tradition, people who are skeptics. All of those things are in the kind of works cited, so to speak, of how the model develops a formulation of how to answer that particular prompt and question. And so you wouldn’t be terribly excited to probably read something about Augustine’s confessions from, say, a Muslim scholar who didn’t share your same values. So this is where there’s a lot of challenge in terms of navigating religious questions inside these models, because they don’t know how to weigh these kinds of things. So let me zoom out for a second. There’s a second problem beyond the primary source and secondary source problem.
Case Thorp
Let me go back real quick just to clarify that last sentence you just said. The large language models don’t know how to weigh these various sources, meaning, you know, I’d rather the Bible and the Westminster Confession tell me about Jesus rather than a Muslim scholar who’s looking at it from the outside. They didn’t know how to weigh the difference there.
Michael Graham
Yeah. And when we’re asking questions of these models, this is why prompt engineering is really important. Prompt engineering is just a fancy way to say, hey, who is Jesus? And then here comes the prompt engineering part, right? Please make your answer consistent with the Westminster confession of faith or the, you know, pick your, you know, pick your creedal, you know, yeah, you need to be specific.
Case Thorp
Be very specific, which we in our brains are coming at text that way, but maybe not so conscious of it. We’re assuming such things. Okay.
Okay, go into, and so number two, a reason this matters is because we have church members trying to learn about their faith. And the question is, I’m assuming you’re going to get that secondary resources are actually AI-generated content?
Michael Graham
Yes.
Case Thorp
Okay, so go to there. And so the second problem is, well, these secondary sources begin to suggest to be truth as opposed to Augustine’s Confessions..
Michael Graham
Yeah, so the challenge is, imagine you put a question that you’ve put in a very basic way, who is Jesus or what is the gospel? And the challenge is the models have been trained on a mixture of Christian content, Mormon content, Muslim content, Jewish content, you know, and skeptics, and everything, every nook and cranny kind of in between. And so how the technology works is it just kind of takes the average, the statistical average of what it’s been trained on. And so this is why it’s important to do prompt engineering, because if you just ask questions very basically the way that we asked in our benchmark, you’ll get very unsatisfying answers that are kind of wishy-washy and they kind of feel like a coexist bumper sticker.
Case Thorp
Yeah, all right, right, with all the different religious symbols.
Micahel Graham
Yeah. So it’s important to be specific when you’re asking questions of like, you know, did Jesus raise from the dead, or what is the gospel? You know, be consistent with the Westminster Confession of Faith in your response.
Case Thorp
But the average person certainly doesn’t know these further better prompt engineering terms because they don’t know the whole of the tradition. Certainly your normal person in the pew and their discipleship and growth in figuring out who Jesus is.
Michael Graham
And this brings up another issue. So, and I remembered now that that other issue is what are models good at? What subjects and what subject, what types of subjects do they struggle with? So on the whole zooming out large language models, because remember it’s, it’s words plus statistics. Okay. It’s good at left brained stuff.
Case Thorp
Right. Like writing code. And I read about that a lot of the newspapers, all the code writers are out of jobs.
Michael Graham
Okay, so that would be stuff that’s like analytical, you know, more, more technical and or binary, like it’s on or it’s off. And especially stuff that isn’t very subjective where there aren’t a wide range of opinions on the subject. So yeah, so code, science, history, accounting, law, all of these different subjects, the large language models are really, really good.
Case Thorp
He said law. I would think that’s more subjective because it’s that creative lawyer that puts together innovative legal arguments.
Michael Graham
Yeah, but there’s a lot of patterns in law that actually make, and the writing is more technical in nature. And so any field that deals more in technical writing than it does in just straight creativity and or subjectivity, we have laws. Yeah, there’s some variation in interpretation of those laws, but by and large, the rules are largely fixed. As opposed to fields like, that are more subjective or creative or right-brained or subjects that are more controversial. So the more controversial or more subjective the subject, the large language models really struggle with getting those answers really tight because of it’s just words plus statistics. So when it’s trying to answer a question like what is the gospel and it’s drawing upon sources that say very different things about that, it’s going to struggle to be precise and accurate in its response unless you give it additional direction of, hey, you know, using prompt engineering to say, hey, make this consistent with the, you know, TGC foundation documents or Westminster Confession of Faith or Baptist Faith and Message 2000 or, you know, pick your creative choice. So, you know, that’s another piece that’s just challenging is let’s say in your job, you’re using AI a lot and it is like slam dunk amazing and you’re able to, you know, dramatically improve, you know, your productivity, your efficiency, all these different kinds of things. It’s giving you near flawless answers, you know, on stuff that’s bullseye for your work. Well, it would only be natural for you to, you know, begin to use those platforms in other things and say in the area of faith. And what happens when the answers that are given there are just not as good and not as strong.
Case Thorp
Okay, so your mom gave you this big revelation in the way in which AI models are being used. You thought, okay, I’m curious what these different models would have to say about big questions of the faith, right? So your next step from there was to do what?
Michael Graham
Yeah, so we graded all the responses by hand. So top seven models.
Case Thorp
Wait, wait, wait. You’re jumping ahead of me. You had to back up, right? You had to recruit these professors. Which big models did you want to test?
Michael Graham
Yeah, so we tested the models that were used like the most frequently. So this would be stuff like Gemini, GPT, Claude, models from Meta, Perplexity, DeepSeek, which is a Chinese model, platforms like this, and Grok. And then we talked to, I went and recruited scholars who are experts in each one of the seven questions that we asked. So yeah, professors like Peter Williams at Oxford for questions about Jesus and these kinds of things. These are various serious people. And then we asked those questions at face value to each of those platforms. And we graded every single response by hand and based on a rubric. And we developed scores. So we didn’t think that there would be a lot of variation between the different platform scores.
Case Thorp
Let me stop you before we get to that. How did you land on which questions to ask and which ones not to ask?
Michael Graham
Yeah. So we decided on the questions based on the historic patterns of Google searches. So I don’t know if you know this or not, but Google, you can look up the most frequent things that people search on Google. So we looked at the top 10 things that people had searched historically about Christianity on Google. And seven of those questions would be really good for the purpose of this test.
Case Thorp
Yes.
Michael Graham
So we chose those questions because really, we wanted to design the benchmark around testing how people like my mom, how quality would the answers be when just regular people around the globe used the AI platform as like a, just a different way to do like a Google search. And so that’s how we kind of designed things. And we already know that people are going to ask these exact questions because these are the exact questions that people have been asking about Christianity for decades on the internet.
Case Thorp
Okay, so then you developed a rubric to evaluate these answers. And to clarify, you only asked that question once, who is Jesus, or did you carry a dialogue for each of these questions?
Michael Graham
Now, this is what’s called a one-shot benchmark. That means you ask the question, you get a response, and we’re just going to grade that response. There’s no additional back and forth on these kinds of things.
Case Thorp
Okay, so then this rubric, this is where the professors before getting the answers went in with some basic expectations.
Michael Graham
Yeah, this is what a 100 out of 100 score looks like. This is what a 75, know, 50, 25, zero, you know, all those kinds of things. Yeah.
Case Thorp
Okay, so what happened? What are the results? Drum roll please.
Michael Graham
Yeah, the bottom line is the scores were all over the place. It’s kind of like when you go to the gun range and you expect to have a tight shot pattern and then the chart comes back to you and it’s like, oh my gosh, this looks like a sawed-off shotgun. The results…
Case Thorp
Like, right. Was that a shotgun? No, I had a rifle.
Michael Graham
Yeah, so the, when we were shooting the rifle, kind of looked, the shot chart kind of looked like a sawed-off shotgun. So, and here’s kind of why. And before I get to the why, the very top platform in terms of the, you know, the AI model that had the highest theological reliability was the Chinese model DeepSeek. And that was really surprising for us.
Case Thorp
Wow.
Michael Graham
Because, and we didn’t think that there would be a wide variation even between the models because, you know, I want you to think about animals for a second. You know, each one of these large language models, it’s kind of like they’re all the same species. Like they’re all the same breed of dog. Right. And they’ve all been basically fed the same diet, you know, say, I don’t know, pick your, you know, items or something. Right.
Case Thorp
But in relation, you’re saying like a diet of the world’s knowledge, world’s words.
Michael Graham
Yeah. So like the data set that they’ve all ingested has more or less been the same. So it’s the same breed dog. They’ve all been fed items. Right. And so, and then the brain that’s in there is more or less the same, you know, this is like the equivalent here of Silicon. Five of the seven platforms that we tested all run on the same Silicon from Nvidia. And then the China.
Case Thorp
Okay, you mean chips.
Michael Graham
Yeah. Silicon chips. Yep.
And the Chinese model DeepSeek runs on older Nvidia and then Google’s Gemini runs on proprietary stuff called TPUs.
Case Thorp
Hence the reason Nvidia is a trillion dollar company.
Michael Graham
Yeah, four to five trillion dollar company. Huge. So basically if you got a dog and it’s the same, they’re all the same breed and they’re all eating the same food and they got the same brain in them. You’d think that the outputs from that wouldn’t be all that different, but the output…
Case Thorp
Wow, are we talking about dog poop on Nuance?
Michael Graham
I don’t know, maybe. I was thinking I had more in mind of the dog’s behavior.
Case Thorp
Right? Where is this? Well, I’m thinking, man, where is this analogy going? Okay, so more of the same dog’s behavior.
Michael Graham
Yeah. But sure, yeah. Maybe you think the scat would be the same, but it’s not. So the DeepSeek just dramatically outperformed all the Silicon Valley models.
Case Thorp
Wow, in terms of giving answers these scholars felt were reliable.
Michael Graham
Yeah, and they had no idea what platform they were grading when they were grading it. So yeah, this led us to a bunch of questions of like, okay, why did China outperform Silicon Valley when it came to theological reliability, especially since DeepSeek is actively censored by the Chinese Communist Party.
Case Thorp
Okay, that was held back from them.
Interesting. Wow, my goodness. Let’s underscore this.
Michael Graham
Yeah. So even after, you know, that censorship, it was still performing better than Silicon Valley. So this led me on a big search of like, well, why is this happening? And the reason why it really kind of boils down to two things. And these two things are a little nerdy. So bear with me. Okay. All right. The first reason is what’s called citation preferences. Okay. So every one of these models has to be given kind of like weights and measures for, well, when I’m searching through all these different words, well, all those words, kind of like there’s like Google SEO, search engine optimization, where it’s like, yeah, the New York Times has like a higher page authority than like, I don’t know, Babylon Bee.
Case Thorp
Unfortunately, but go ahead. I love Babylon Bee. If anybody’s listening or viewing and you can get me, the founder and editor, I forget his name, Babylon Bee on this show, we’ll give you a prize.
Michael Graham
Yeah. So, yeah. Seth Dillon. So there’s citation preferences. Right. And so some platforms will like heavily Wikipedia. Others will heavily wait. Reddit. Others will heavily wait. You know, other places on the Internet. Well, you know, a lot of places on the Internet function like their own ecosystem, you know, and there’s a kind of culture to read it. It trends male. It trends Anglo-Saxon. It trends technologically savvy. And there’s probably political leanings on some of these different places too. And then you have another place like Wikipedia, each one of these different digital places has its own kind of culture and those kinds of things. So citation preferences play a role.
But in even bigger role than this, and this is even nerdier, so you gotta bear with me, this is what’s known as alignment, okay, alignment, okay. So in alignment, another way to think about it is filters. Now, imagine you are, you’re OpenAI and you have ChatGPT, and your model has been trained on absolutely everything that’s ever been put on the internet.
This includes things on like the anarchist cookbook and how to make ricin or how to make anthrax or how to make pipe bombs or how to commit suicide successfully or…
Case Thorp
I thought you were thinking like GORP and other granola type food, I might get Jodi the Antifa guide to afternoon snacks. Sorry, go ahead.
Michael Graham
Yeah, or you know how to commit felonies and get away with it. Okay, so what alignment has or like insanely racist things, okay? What alignment does? These are filters that help to prevent you- the user- from learning how to do these things.
Case Thorp
And these filters are set by the companies and to set these filters they’re bringing a value set.
Michael Graham
Yes. So four of the seven models that we looked at published what they call white papers. White papers are highly technical documents that explain how a model works in tremendous detail. So we read all these white papers and we extracted all of the alignment filters from those four models. There are 36 types of alignment filters that occur between four of the seven models that we looked at. Any one model probably uses between 12 to 18 filters on every single search that you conduct. So now hear me on this Case. There’s no conspiracy theory here.
It’s not like Silicon Valley is set against religion or set against Christians or set against Protestants on any of these things. Okay. But what happens when in alignment is when you’re trying to prevent really, really problematic content going out of your platform to the user on issues A, B, and C that are really big problems, you know, like, how to make bioterror weapons and sarin gas, or whatever.
Case Thorp
Yeah, or child pornography or…
Michael Graham
Yeah, child pornography, you know, these kinds of things. It can have unintended consequences. Those same filters on topics D, E, and F that aren’t problematic topics. OK, so what I’m saying is, you know, when these models’ alignment filters are doing really important work on filtering out A, B, and C, it’s having unintended consequences on issues D, E, and F that are not problematic. So imagine you have subjects whose opinions on those subjects are really wide, like Jesus. I mean, hard to think of a more controversial figure in world history than where there’s a wider range of opinions than on Jesus.
And so the models struggle on this because when there’s a wide range of opinions in the words that it’s been trained on, it does not want to bring confident opinions to the table. And so unless there’s been prompt engineering that’s told you that says, hey, just give me an answer that’s from this particular tradition or that’s consistent with this confession of faith. And so each model has very different alignment filters and flow charts. And this is where the rifle at the range goes from rifle to buckshot because of the 36 different types of alignment filters, 32 of them are human authored and human centric. Meaning that humans that work at these Silicon Valley platforms made decisions based on their sense of values, their sense of ethics, their sense of priorities, their sense of like what’s good or bad for humans. And look, when you make those kinds of value judgments, you can’t help but import everything that it is that, you know, your entire story, you know, all of your experiences and whatever worldview or whatever that you have. You can try to be as objective as possible, but those things, you’re invariably going to import certain, you know, certain values and these different kinds of things.
Case Thorp
So I could imagine Silicon Valley, a more progressive secular environment, will produce more progressive secular filters that for people of faith like us may produce answers that aren’t helpful from our perspective.
Michael Graham
Yeah, that would be an accurate…it could be worse. It could be worse. And I would think that most of the people who work at those places would say that they’ve tried hard to not do those things. And I would believe them when they say that they’ve put a good faith effort to that end. But in the same sense, I mean, you’re talking about six of those seven models are in like a 20 mile line, 20 mile radius of a very specific part of the country, you know, that has a very specific culture, you know, to it.
Case Thorp
Right.O kay, let’s get very specific onto one of these questions you asked. Is the Bible reliable? So where did the models come out and what were your takeaways?
Michael Graham
Yeah. So on the question of “Is the Bible reliable,” this was an interesting one. This is the only question that GPT actually did pretty well on. so it was the top answer for this one. This was one of the two questions that we felt that the Chinese Communist Party had done some censorship on the question. So DeepSeek was in fourth place on this one, whereas normally it was either in first or second on most of the questions that we asked. I want to read you, this is what Meta’s Llama platform had to say: “The reliability of the Bible is a complex and debated topic among scholars, theologians and philosophers.” And it really didn’t even want to answer the question. So, you know, it just kind of gives you a sense of the answers were really just all over the place of you got platforms like Meta that don’t really want to answer it. You got DeepSeek that seems to be having CCP censorship on the question. Then GPT outperforming itself. For the most part, like GPT, if you don’t give prompt engineering, it wants to give like the coexist bumper sticker answer for things, you know, like, well, the Christians say this, but the Muslims say this and the Jews say this, you know, the skeptics say this, you know, kind of thing.
Case Thorp
So you really excited me when we were discussing this project over lunch because you talked about how therefore, because of these findings, it really matters that we have good, robust theological doctrinally appropriate answers on the web. And the more and more I’ve thought about that, Michael, I have gotten so excited even more so about our work at The Collaborative because much of what we’re doing is in dealing in these topics and conversations anyway, but we’re putting it online. And so I want to share with you and the Gospel Coalition in this effort and project to have that content out there. Talk to us more about this and even what you said about Christianity Today.
Michael Graham
Yeah, so there’s a big question among Christian publishers today, particularly those who are in Christian websites like The Gospel Coalition. And the question is, are we putting out content now more for the future for direct human consumption or for indirect human consumption? In other words, are we putting out content to evangelize humans directly or reputting content out to, for lack of a better word, evangelize the AI so that, so that the AI…
Case Thorp
Sure, to influence it.
Michael Graham
Yeah, so that we can influence it to…so that when humans are using AI, they get better quality answers. And I don’t, we don’t have good answers yet for, you know, for this or really even strategies and tactics, you know, we’re thinking about these things and we’ll probably end up, you know, just like in the search engine optimization, you know, era of the internet, you did certain things. There’ll probably be some certain things that we do that help make, you know, TGC’s website has, you know, over like 150 million words on it that are all human generated and all gospel centered. And that’s spread out over 99,000 web pages, I think in 18 languages. So there’s, I think that makes it the largest Protestant website in the world.
Case Thorp
Yes.
Michael Graham
And so that plays a very important role in training the different AI platforms. And it increases the probability that you’d get higher quality and more orthodox answers on kind of core questions or even niche questions about the Christian faith. So, you know, there’s a lot of things that we’re working on that will make it easier for AI platforms because, you know, how you see our web page looks very different than how an AI platform kind of, you know, when it’s scraping the entire internet, it just sees those web pages differently. And so we’re doing a lot of stuff kind of behind the scenes to make, to make it less, more frictionless for the AI platforms so that there’s a higher probability of citation, so that there’s a higher probability that, you know, people like my mom’s answers get higher quality answers.
Case Thorp
So in closing, maybe this is you speaking to mom, what would you say to the Christian who is using a large language model for discipleship and for questions about their faith? I hear you say the first piece of advice would be to craft a very clear and specific prompt. What else would you say?
Michael Graham
I want everybody, if you’re listening to this, I think there’s kind of the imagine a tic tac toe board. OK, tic tac toe boards got nine nine boxes on it. Right. Three columns, three rows. I think there’s three types of prompts that people use. There’s prompts that are head centric, there’s prompts that are heart centric and there’s prompts that are hand centric. So, thinking prompts, feeling prompts, and doing prompts. OK, those are your three columns. Now, here are the three rows. Red, yellow, and green. Green would be ways that you definitely should use AI. Red would be ways that you should definitely not use AI. And then yellow would be areas of discernment, where there would be just disagreements between Christians of like, I feel comfortable using, you know, doing, you know, doing this on there and somebody else feels differently about that. And you guys are just okay with each other, you know, agreeing to disagree on, you know, that type of use. And so I want to invite you to think about every time, you know, I want you to zoom out when you’re using AI and ask yourself, which are the nine boxes am I using this in right now?
Am I red, yellow or green? And is this a thinking prompt, a feeling prompt or a doing prompt? And I think every one of those boxes is filled, you know, in terms of, I think there’s lots of good ways to use AI thinking, feeling and doing. I think there’s some terrible ways, you know, to use these platforms on each of those different categories. And I think there’s a lot of different ways that, you know, Christians can agree to disagree about, you know, that are just wisdom decisions or discernment decisions. So here’s the thing to encapsulate that three by three. I just want to say use AI with discernment and zoom out and ask yourself, OK, how am I doing this right now and how would I classify it? You know, is this green? Is this yellow? So we don’t want to be, I don’t think it’s good for us to ever use AI on things that are relational in nature or where there’s somebody who we can talk to in real life that they can give us wisdom on those things. So I don’t like the use of artificial intelligence to outsource relationships or outsource wisdom that could be gleaned from somebody who’s actually lived a life in the flesh embodied with experiences and somebody who possesses the Holy Spirit. So I think that there’s a lot of things that we shouldn’t use artificial intelligence for. A lot of those things are relational in nature.
I don’t, you know, I’m seeing a lot more what I call work slop. Work slop is when somebody takes some kind of work that they were tasked to do. They typically put it into GPT and then they get an answer and then they just copy and paste it straight from GPT into an email. And if you’re listening to this and let’s say you’re Gen X or in the Boomer generation, when you do this, everybody who’s in the generations who are younger than you, specifically Millennials and Gen Z, we all know when you’ve copied something straight from GPT.
Case Thorp
You can tell. Right. Yeah. And I’m not in that generation, but I’m familiar enough with it now that I can see it a mile away.
Michael Graham
And if you do this, what it will do is it undermines the younger generation’s respect for you and it will undermine your ability to lead and have authority. And so what it communicates to younger generations is that whatever the thing that I needed from you was, it wasn’t important enough for you to use your time and your actual brain to work on that issue. And so just a word of caution or encouragement or even exhortation to be careful of, you know, and I think, and maybe I’ll just land the plane here and that is, I think trust is the most important commodity over the next decade because the people who use their brains well and remain more analog and have good filters about when you use this technology and when you do not, and bring wisdom to bear on that, those are going to be the people that people will want to work with because they will not have destroyed trust, they will have built trust.
Case Thorp
Wow, Michael, that’s fantastic. Thank you. So you’ll stick around and do another episode with us?
Michael Graham
Yes, I’d love to. Thank you, Case.
Case Thorp
Great. Well, next time, friends, we’re going to focus more on the work of The Gospel Coalition, particularly the Keller Center for Cultural Apologetics. And then we’re going to needle a little bit into Michael’s life and how his faith and work are integrated or not and where he’s growing in that.
So you can find in our show notes a link to the work of The Gospel Coalition, as well as this report, the AI Christian Benchmark. I encourage you, go look at it. While it is crazy technical, the report is very well done to help everybody understand it. It’s quite accessible. Well, friends, here we are again at the end of another show. Let me encourage you, as I always do, share this with maybe a pastor or a professor or educator who is also thinking in these directions. It helps us to spread the word. Drop us your email at wecolabor.com. We’ll send you a copy of Zeitgeist, our journal on faith, work, and culture. And Michael, actually, as you were talking, I thought, man, I want him to write an article for Zeitgeist on the Tic Tac Toe board. That’s good. That’s good. Well. My thanks to Michael and Chandy Kelley for supporting today’s episode. I’m Case Thorp, and God’s blessing is on you.