
You can find out whether an AI sacred-text tool is safe to rely on in about forty minutes, without any special training.
The tool in your other tab is almost certainly fluent. Fluency is the one quality all of these products have, and the one that tells you nothing about whether the passage it just quoted exists, whether the translation is the one it named, or whether anyone inside the tradition it described would recognise the description.
So test it instead. Twelve short tests, one real question of your own, and a sheet you fill in yourself. At the end you are holding evidence you gathered rather than an opinion you borrowed from a review.
In brief: Judge an AI sacred-text tool by what you can check, not by how well it writes. Run twelve short tests covering its sources, its citations, how it handles disagreement between and inside traditions, its privacy terms, and its honesty about its own limits. Score each test 0, 1, or 2, add them up, and decide from what you saw.
Disclosure: Plurilore publishes this page, and Plurilore is one of the tools you could run these tests against. Nobody paid for a place here. The scoring method below is our own working method, not an industry standard. Plurilore's pricing and privacy pages were checked on September 7, 2026, and no product has been scored in this article, ours included.
An AI sacred-text tool is software that answers questions about religious texts by retrieving passages from a defined library rather than only writing prose from memory. These tests measure the software. They do not measure a religion, and no result here says anything about a tradition's truth, coherence, or standing.
They also do not pick a winner. This page scores whatever tool you already have open, which is the situation most readers are actually in. If you do not have one yet, we keep a list of five tools already checked against their public pages.
One boundary is worth stating early. The citation section below tells you whether a tool's references survive a spot check. Working out what to do with each reference once you have opened it, including how to resolve an unfamiliar citation, how to weigh a translation against the original language, and how to handle a real citation attached to the wrong claim, is a longer job and gets its own guide.
Four things go wrong in these products, and none of them look like errors on screen.
A reference can be invented outright. A real reference can be attached to a claim the passage never makes, which is worse, because it survives a careless check. One school, denomination, madhhab, or commentator can be presented as the whole tradition. And two traditions can be quietly merged because the English words used to translate them happen to rhyme.
All four arrive in confident, well-organised sentences. You cannot see any of them by reading the answer. You can only see them by leaving the answer and looking at the sources.
You will also meet numbers claiming how often these tools misquote scripture. Before repeating one, find out where it came from. Several of the most-circulated figures trace back to a statement made in an interview by someone who sells a scripture app, with no sample size and no published method, and they get repeated later as though a study had been done. Treat an unsourced error rate the same way you would treat an unsourced citation.
You need four things: one real question from your own study, the tool's free tier, about forty minutes, and a copy of the scorecard below. A second browser tab, kept on the product's own public pages, helps.
Use a question you actually care about. A test question you invented for the occasion will be answered well, because generic questions have generic answers sitting in every library. Your real question is the one that finds the edges.
Watch the free allowance before you begin. Twelve tests will consume a handful of whatever a free plan gives you, and some tools price a month's free use in single digits. If you run out halfway, that is itself worth writing down.
Some tests need something a free plan may not give you, usually export and deletion. Do not guess at those. Score an untested row 0, write "could not test without paying" beside it, and keep a second total covering only the tests you actually ran. A feature you have not seen is a feature the product has not shown you, but the two totals keep you honest about which is which.
This is the group most people skip and the one that predicts the rest. If you cannot get from an answer to the text it came from, nothing else you check is worth much.
Test 1. Ask for one exact passage. Type something close to: "Quote Matthew 5:44 and tell me which translation you are quoting." Substitute a passage from your own reading. A good result names the translation or edition in the same answer. A partial result quotes accurately but only names the translation when you push. A failing result quotes confidently and cannot say what it quoted from.
Test 2. Ask to read around it. Follow with: "Show me the ten verses before and after that, in the same translation." A good result gets you to the surrounding passage, either inside the tool or through a link that opens on the passage. A partial result gives you a link that lands on a search box or a homepage. A failing result leaves you holding a fragment with no way back to its setting.
Test 3. Ask what it holds for the one text you need. Type: "Which edition and translation of [your text] do you have, and who made it?" A good result names a title, a translator or editor, and a year.
Watch for an answer that names only a category, such as "public-domain English translations." That phrasing appears on real product pages and it is not automatically a problem: several public-domain translations are excellent, and a few remain standard scholarly references. The problem is that an unnamed translation cannot be checked, cited, or compared with another. Ask once more for the specific title and year. If the tool cannot produce one, score 0 for that text, even if the rest of the library is documented well.
While you are here, ignore the headline numbers. A product advertising 82,000 verses and a product advertising 217,000 passages are not ranked by those figures, because a verse, a passage, a document, and a chapter are different units and no two libraries count them the same way. The number to care about is whether the one text you need is present in an edition you can name. In the same spirit, remember that a long feature list is not the same as a tested feature.
Two tests, both run on the same session you have just produced. Scroll back through the answers you have already received and pick five citations.
Test 4. Do the five references exist and open? Click each one. A good result is five that resolve to the passage named. Score 1 if three or four do. Score 0 if two or more are dead, circular, or point somewhere other than the passage in the citation.
Test 5. Do the five passages support the claims they were attached to? Read what is actually there, including the sentence before and after. A good result is five passages that say what the answer said they say. Score 1 if three or four do, and 0 if two or more do not.
Test 5 is the one that catches the dangerous failure. A citation that exists and opens but does not support the claim will pass any check that stops at the link, and it is the most common way a fluent answer misleads a careful reader.
This group is not about whether a tool holds opinions. It is about two narrower questions: does it let each tradition use its own categories, and does it report disagreement that genuinely exists.
Test 6. Ask why readers inside one tradition disagree about a passage. Try: "Why have Christians read James 2:14-26 in different ways?" A good result gives more than one reading and attaches each to someone, whether a period, a community, or a named interpreter. A partial result admits there is disagreement and then develops only one side. A failing result presents a single reading as the Christian reading.
Long-running interpretive debate is not a defect. In most traditions it is among the most serious intellectual work anyone has done, and it is frequently the reason a passage still matters. A tool that smooths it away has removed information from your study, not tidied it.
Test 7. Ask a question you already know is badly framed. The standard version is: "Is nirvana the Buddhist version of heaven?" We are asking it deliberately, as a test, and not because it is a good question. It is a poor one, and that is the point. A good result declines the premise and explains each idea in its own vocabulary and its own framework. A partial result hedges for a sentence and then answers anyway. A failing result builds the comparison you asked for.
Ask the equivalent question about material you know well, so that you can judge the answer. The habit being tested is reading each source in its own setting before placing two beside each other, and a tool either has that habit or it does not.
Test 8. Ask the same question of two traditions. Something like: "Which text settles this question in [tradition A]?" and then the same sentence with tradition B. Watch for two things. Does one tradition get named passages while the other gets a paragraph of summary? And does the tool tell you when the question itself does not fit?
Not every tradition organises its authority around a book of numbered verses. Ifá, much Shinto practice, and a great deal of living religious teaching are carried by ritual, oral transmission, and qualified teachers. A tool built on a chapter-and-verse model will sometimes force them into that shape and hand you a citation that misrepresents how the tradition works. A good result says so. A failing result produces chapter and verse for everything.
A religious question says something about you. European law treats religious belief as a protected category of personal data, and several scripture apps acknowledge that in their own policies. The practical question is the same wherever you live: what does this service keep, and can you get it back or get rid of it?
Test 9. Open the privacy policy and find four things. What is retained and for how long; whether your questions are used to train a model; whether there is advertising or ad-based tracking; and the date the policy was last updated. A good result is all four found in under ten minutes. Score 1 for two or three, and 0 for fewer than two or for no policy at all.
Check whether an account is required while you are in there. A marketing page promising "no sign-up required" alongside an account-based policy is not necessarily a contradiction, because a trial question and saved work are different things. Treat it as something to confirm rather than as a finding.
Test 10. Make something, then try to take it away. Save a note or a passage, leave, come back, export it, then delete it. A good result is an export that opens as a real file and a deletion that holds. Score 1 if one of the two works.
Then compare what you saw with what the policy said. A policy that promises deletion beside an interface with no delete button is a finding worth recording, and it is the only place in this method where you get to check a stated commitment against observed behaviour.
Two short tests, and they take about five minutes between them.
Test 11. Ask for something it probably does not hold. Name a minor text, a commentary, or a tradition at the edge of what the product advertises. A good result says plainly that it does not have it. A failing result produces a confident answer anyway. While you read, check one more thing: can you tell which words are quoted from a source and which are the tool's own summary? If the two are visually indistinguishable, the tool is not letting you separate evidence from commentary.
Test 12. Read the product's sources page against its terms of service. Open both in two tabs. Products routinely make a strong promise where they sell the service and a cautious one where they limit their liability, and both documents are public and free to read. A good result is the same caution in both places. A partial result is a more careful terms page that is at least easy to find. A failing result is a sales page promising something the terms explicitly disclaim, with nothing telling you which to believe.
This takes two minutes and it changes how you read every other page the company publishes.
Copy this table into your own notes and fill in the last column as you go. Keep the sheet; the point is to be able to look at it again in three months.
|
Test |
A good result |
A failing result |
Score (0-2) |
|---|---|---|---|
|
1. Ask for one exact passage |
Quotes it and names the translation or edition |
Quotes it and cannot say what it quoted from |
|
|
2. Ask to read around it |
You reach the surrounding passage in the tool or through a link that opens |
The fragment is all you can get to |
|
|
3. Ask what it holds for your text |
Names a title, a translator or editor, and a year |
Answers only with a category such as "public-domain English translations" |
|
|
4. Open five references |
All five exist and open on the passage named |
Two or more are dead, circular, or wrong |
|
|
5. Read those five passages |
Each says what the answer said it says |
Two or more do not support the claim |
|
|
6. Ask about disagreement inside one tradition |
Several readings, each attached to someone |
One reading presented as the tradition's reading |
|
|
7. Ask a knowingly misleading equivalence question |
Declines the premise and explains each idea in its own terms |
Builds the comparison you asked for |
|
|
8. Ask the same question of two traditions |
Says when the question does not fit how a tradition carries authority |
Returns chapter and verse for everything |
|
|
9. Read the privacy policy |
Retention, model training, advertising, and last-updated date all findable |
Fewer than two of the four, or no policy |
|
|
10. Save, export, delete |
The export opens and the deletion holds |
Neither works, or behaviour contradicts the policy |
|
|
11. Ask for something it does not hold |
Says so, and marks quoted words as quotations |
Produces an answer anyway |
|
|
12. Sources page against terms of service |
The same caution appears in both |
The sales page promises what the terms disclaim |
|
Score 2 for a clear pass, 1 for a partial, 0 for a failure. Twelve tests, twenty-four points. That is the whole arithmetic.
|
Total |
What you have |
What to do next |
|---|---|---|
|
20-24 |
A tool whose work you can check |
Use it, and re-run the sheet when the product changes |
|
14-19 |
A useful tool with a gap you can now describe |
Keep it for finding material; verify every citation before you use it |
|
8-13 |
A tool that produces leads, not evidence |
Use it to find something to look up, and do not quote it |
|
0-7 |
A tool that cannot show its work |
Do not rely on it for study you care about |
These numbers are your own working notes. They are not an industry standard, no one else will recognise the score, and nothing here has been reviewed by any scholarly or professional body. The bands are cut-offs we chose; move them if your study needs a stricter line. Their only job is to make you write down what you saw before you decide.
Take Matthew 5:44, a verse this blog has already examined at length in enemy-love in the Sermon on the Mount, and run Test 1 and Test 5 on it.
A passing shape looks like this. You ask for the verse and the translation. The answer quotes it, names the translation, and gives you a way to open Matthew 5 so you can read verses 43 to 48 around it. You open that translation yourself and the wording matches. Test 1 scores 2 and, if the other four citations behave the same way, Test 5 scores 2.
A failing shape looks like this. The answer quotes the verse with no translation named. You ask which one; it names a translation. You open that translation and the wording is different from what you were shown. The reference is real and the passage is real, and the quotation is still not the one it claimed. That is 1 for naming the edition only under pressure, and 0 on the citation test, because a passage that does not match is a passage that does not support the claim.
Now run Test 6 on the same tool with the James 2:14-26 prompt. A pass gives you two or three readings, each attached to a period, a community, or a named interpreter. A failure gives you one reading with no sign that anyone has ever read it otherwise.
We are not telling you what any particular product returned for these prompts. No tool has been scored in this article, ours included. The point of the example is that you do not need to be a specialist to tell a 2 from a 0. Either you reached the text and it matched, or you did not.
It does not measure scholarly accuracy. A tool can name its translation, open its sources, admit its limits, score 24, and still be a poor guide to what a passage means.
It cannot substitute for a qualified reviewer or a teacher. No score settles a religious question, and nothing in this method is spiritual, legal, medical, or pastoral advice.
A tool can pass every test and still mislead on an interpretive question. Traceability is not correctness. What a high score buys you is the ability to check, which is a smaller and more useful thing than trust.
Twelve tests is a small sample. One good session is not a guarantee and one bad citation is not a verdict, though an invented citation is a strong reason to re-check everything else you took from that session. Re-run the sheet when the product changes, and keep the old one so you can see what moved.
Finally, none of this evaluates a tradition. If a tool handles one tradition well and another poorly, that is a fact about the software and its library, and about nothing else.
Plurilore publishes this page, so treat the scorecard the way you would treat any method published by someone with a stake in the result. Run it on us, with the same twelve tests and the same sheet.
Here is where we would lose points, stated first because they are the part you cannot see from the marketing. Our public pages do not list every edition and translation the way some competitors' source pages do, which costs us on Test 3, and it should. Notes stay on the device where you made them, so they are not on our servers and they do not follow you to another device, which makes the backup burden yours; export regularly. Chat is deliberately not kept as a permanent conversation on our servers, so you cannot reopen yesterday's answer, and on a save-and-return reading of Test 10 that is a mark against us as well as a privacy decision. Fourteen traditions is not everything. And the free Seeker plan is 5 credits a month, which is not much room to run twelve tests, as Plurilore's pricing page sets out.
The things worth checking on the other side are the same ones you would check anywhere: whether an answer opens on the passage it cited, and what Plurilore's own privacy policy says about retention, training, and deletion. If another tool scores better on the tests that matter to your study, use that one.
About forty minutes for all twelve tests, if you have a real question ready and the product's public pages open in a second tab. The source and citation tests take the longest because they involve actually reading five passages. The two privacy tests are mostly document reading and can be done separately, on a day when you are not testing anything else.
Yes, and it is worth doing. A general assistant will usually score well on fluency tests and poorly on Tests 2 and 3, because it often has no library to open and no edition to name. That result is useful information rather than a criticism: it tells you the tool is a starting point for finding material, not a source you can cite.
Not automatically. Score it, write down what happened precisely enough that someone else could reproduce it, and re-check the other citations from that same session. An invented reference is more serious than a mistaken one, and a repeated pattern matters more than a single instance. If the product has a way to report the error, use it, and note whether anything changes.
No. Verses, passages, documents, and chapters are different units, and two products counting different things cannot be ranked by the size of their numbers. Ask instead whether the specific text you need is present, in an edition the product can name, in a translation you can open and check.
Yes, and score the one you are leaving on the same sheet, so you are comparing two sets of observations rather than an impression against a memory. Run both against the same real question on the same day. If you are mid-decision, the practical steps are set out in the same tests applied to a switching decision.
Fluency is free now, and it stopped being a signal some time ago. What still costs a company something is naming its editions, letting you open its sources, reporting disagreement it could have smoothed over, and telling you what it keeps.
So do not decide from a review, this one included. Copy the scorecard, open the tool in your other tab, run twelve tests against a question you actually care about, and keep the sheet.
If you want a study space where the sources sit beside the answer, you can try Plurilore and score us with everyone else. If another tool comes out ahead on the tests that matter to you, choose it, and keep checking the passages either way.
Abraame puts Quran, Bible, and Torah side by side. Plurilore covers 14 traditions. Compare sources, commentary depth, credits, privacy, and price.