Skip to main content

Product × Engineering case study

We Didn’t Fix JainKosh’s Hallucinations With a Bigger Model

A product case study in turning fluent-but-unsupported answers into a source-grounded, human-reviewed knowledge experience.

JainKosh product team11 min read

Problem

Fluent answers could outrun their evidence—and one polished mistake could erode trust in the entire product.

Goal

Make unsupported output harder to produce, uncertainty visible, and every confirmed error useful to correct.

Outcome

A bounded source and review system; the curated July hardening bank reached 120/120, without claiming universal accuracy.

Live product capture
JainKosh AI chatbot on JainKosh.org in English mode showing a karma philosophy question and a detailed answer about karma as physical matter, with headings, bullets, and the language selector set to English.

Product evidence · JainKosh.org chatbot

The chatbot lives on the main JainKosh website, not a separate app

An English question about karma philosophy answered inside the JainKosh.org widget. The language selector, question, and structured answer are all visible. Capture from production on August 24, 2026; no account or personal data is shown.

The first JainKosh chatbot had a product problem disguised as a model problem.

It could answer in Hindi, Hinglish, and English. It streamed quickly. It sounded confident. But confidence was exactly what made a wrong verse, a mixed tradition, or an unrelated source link dangerous: the answer looked finished before it had earned trust.

For a community knowledge product, that is not a minor quality bug. One polished mistake can make readers question every answer that follows.

So we stopped asking, “Which model hallucinates less?” We asked a more useful product question:

How might we make an unsupported answer difficult to produce, easy to detect, and useful to correct?

That reframing changed the roadmap. The work was no longer a model upgrade. It became a product system: source boundaries, risk-based routing, final verification, visible feedback, volunteer review, and regression tests.

01. The product problem: confidence outran evidence

The early architecture was intentionally small: a JavaScript widget on JainKosh.org called a Next.js API, which sent a language prompt to a general model and streamed the result.

That MVP proved something important: people could ask natural questions without learning MediaWiki navigation. It also exposed three user risks.

  • A plausible answer could cite the wrong evidence. One response reused the wrong illustration for a Moolachar verse; another blurred two sections of Chhahdhala.
  • A general answer could cross a doctrinal boundary. Digambar-default responses sometimes introduced other-tradition details without the user asking for a comparison.
  • A reported mistake could return. Prompt edits improved wording, but they did not create a durable invariant across paraphrases, models, or web and mobile routes.

The prose was polished. The evidence chain was not.

Calling all three failures “hallucination” hid the product decisions we needed to make. Retrieval failures, tradition-policy failures, unsupported citations, generation drift, route inconsistency, internal-text leakage, and broken feedback each belonged to a different layer.

02. The goal: design for recoverable trust

“Zero hallucination” was never a credible product promise. Our goal became narrower and measurable:

  1. Ground source-critical answers in JainKosh material.
  2. Block known high-risk errors before they reach a reader.
  3. Show uncertainty instead of inventing precision.
  4. Keep the experience consistent across web, iOS, and Android.
  5. Turn confirmed mistakes into tests so the same failure is harder to repeat.

This is the difference between an output and an outcome. Shipping retrieval, a review queue, or a new model is an output. The outcome is a reader getting an answer that is more supportable—and a volunteer being able to repair the system when it is not.

Architecture decision

From prompt-only answers to a layered system

The model became one component in a larger evidence and review path.

Early MVP

Question
Prompt
Model
Answer

Fluent output, no enforced evidence boundary

Current path

Question
Policy + evidence
Model draft
Verify + display
Answer
Human review + monitoring

03. Four product decisions changed the system

Decision 1: give the model a source boundary

We built an offline ingestion pipeline around JainKosh’s MediaWiki pages and uploaded study material. Each source keeps its title, URL, revision identity, checksum, timestamp, and section path. Searchable chunks preserve language and quality signals.

At this review point, the production corpus contains 1,277 active sources and 27,733 chunks.

The live retrieval path is deliberately practical:

  • prefer a strong title match when a user names a granth or JainKosh page;
  • run ranked lexical retrieval across normalized Hindi, Hinglish, Romanized, and Jain terms;
  • remove weak, off-topic, navigation, test, and sample pages;
  • send the model a compact packet of text, title, section, URL, and quality metadata.

Retrieval supplies evidence. It does not prove that the draft used every supplied passage correctly.

Decision 2: route by risk, not by one universal prompt

Some questions can use open generation. Others should not.

JainKosh now resolves every question into a typed answer policy. That policy decides whether to use a direct template, guarded generation, source-critical generation, or open generation; whether retrieval and a source line are required; and whether a tradition label or explicit comparison is needed.

Known doctrinal landmines have canonical guards: aliases, prohibited claims, required context, verified source references, and language-specific fallback behavior. A small set of stable, high-risk facts can bypass model improvisation entirely.

An OKF-style concept layer helps recognize aliases, related ideas, and “do not confuse with” boundaries. It guides routing; it is not treated as proof or exposed as a citation.

Decision 3: verify before display

The model’s draft is not automatically the user’s answer.

Before display, JainKosh can reject a prohibited claim, replace an unsafe response with a canonical fallback, prevent unsolicited tradition mixing, remove internal routing text, strip unapproved URLs, flag unsupported exact citations, or say that a source still needs verification.

After internal planning text once reached the interface, we also stopped exposing low-risk answers token by token. The provider can stream internally, but the product buffers the draft, applies the final policy, and only then shows the answer.

Decision 4: keep one answer path across products

Web, iOS, and Android once risked becoming three different AI products. We moved accuracy logic into one shared answer planner and native chat handler. The apps remain native, while policy, retrieval, verification, citations, and answer capture are centralized.

That means one verified correction can improve every surface without waiting for three independent implementations.

Architecture decision

The anti-hallucination defense stack

Different failures are stopped at different layers; no single layer claims perfect accuracy.

1. Input normalizationHelps prevent: Spelling and language drift
2. Answer policyHelps prevent: Wrong route or evidence mode
3. Canonical guardsHelps prevent: Known doctrinal landmines
4. Concept routerHelps prevent: Related ideas collapsed together
5. Retrieved evidenceHelps prevent: General-training improvisation
6. Model draftHelps prevent: Language-aware synthesis
7. Canonical verifierHelps prevent: Prohibited claims and recurrences
8. Citation + display policyHelps prevent: Unsupported links and internal planning
User-facing answer

04. The interface became part of correctness

A trustworthy answer is more than its factual paragraph. Readers also judge the language, formatting, source reference, waiting state, follow-up prompts, feedback controls, and whether the same question behaves consistently next time.

The chatbot lives on JainKosh.org itself—not behind a separate login or app download. Open the widget, pick a language, and ask:

  • language modes are explicit, with English, Hindi, and Hinglish selectable directly in the widget;
  • long explanations use headings, bullets, and bold terms;
  • source-critical answers can name a granth reference;
  • follow-up suggestions are phrased as useful next questions;
  • quality monitoring has a direct on/off control;
  • the same architecture powers desktop and mobile with consistent answer policy.
Live product capture
JainKosh AI chatbot on JainKosh.org mobile, English mode, showing a karma-and-soul question with answer, monitoring toggle on, and Send button.

Product evidence · Mobile experience

The same product works on mobile with the same English answer quality

Mobile view of the same chatbot on JainKosh.org. The question, answer, language selector, monitoring notice, and input are all present. Capture from production on August 24, 2026.

The screenshot is product evidence, not a perfection claim. The mobile view shows the same English answer quality as desktop, but the current system does not yet prove at sentence level which retrieved passage supported every phrase. That gap is on the roadmap.

The same answer policy reaches real phones

The product is not confined to a website. A native iOS app ships to testers through TestFlight, and an Android app ships through Google Play Internal Testing—both using the identical retrieval, verification, citation, and review infrastructure.

The native apps are not wrapped web views. They use SwiftUI and native Android components, with onboarding, a Today daily-practice tab, Jain Calendar with color-coded tithis, full-text search, and an AI tab that calls the same shared answer planner as the web widget. Vikas Ji, Arpit Bhaiya, Shilpy, and several family testers have been testing builds since July.

Live product capture
JainKosh native iOS app onboarding with English language and Deep answer style selected, showing Dictionary, Granths, Bhajans, Lectures, and Panchang content categories.

Product evidence · Native iOS app

The same product also ships as a polished native iPhone app via TestFlight

Native iOS onboarding with English and Deep selected. The same retrieval, verification, and answer policy that runs on JainKosh.org powers the native app. QA capture from iPhone 17 Pro simulator; debug-only footer excluded.

The screenshot shows the English onboarding flow. Every answer that follows—regardless of platform—travels through the same retrieval, policy routing, verification, and feedback capture. One verified correction can improve iOS, Android, and web without waiting for three independent changes.

05. Human review closed the loop

Feedback that disappears into an inbox does not improve a knowledge product.

JainKosh captures the question, final answer, platform, language, model context, source cards, and user response when quality monitoring is enabled. QA and internal activity are marked separately so test traffic does not masquerade as real usage.

A volunteer review queue prioritizes reported answers, negative feedback, incorrect-answer signals, and repeated questions. Reviewers can mark an answer verified, incorrect, needing discussion, or spam; add a category and note; and leave an audit record. “Answer looks correct” is an intentional resolution—review must be able to clear a false alarm as well as confirm one.

At this review point, the system holds 2,062 production answer records and 227 recorded review actions. These are queue records, not user counts.

The final step is the most important: a confirmed correction becomes a regression fixture. The fix is not complete when a prompt changes. It is complete when the relevant router, verifier, citation, parity, or UI test would fail if the bug returned.

Architecture decision

The feedback flywheel has a human gate

Automation helps diagnose and test. A reviewer still decides what should change.

1Question
2Answer + evidence
3User feedback
4Review queue
5Human decision
6Agent diagnosis
7Test + fix
8Verified release

Evaluation sends new failures back into the next review cycle.

06. We measured complexity—then removed it

A PM’s job is not to approve the most advanced architecture. It is to keep complexity only when it improves the user outcome.

ChoiceWhat the evidence showedProduct decision
A bigger model as the primary fixModel changes did not create a source boundary or prevent known repeat errorsRejected as the strategy
Vector retrieval in productionBetter Romanized recall in one probe, but higher storage, timeouts, and weaker full-answer scoresRemoved after a reversible trial
Unverified token streamingFaster first text, but unsafe or internal text could appear before checksBuffer, verify, then display
Human review without regression testsA reviewer could correct one answer while the same class of error returnedConvert confirmed errors into contracts

The vector experiment is the clearest example. The first implementation pushed database storage to 718 MB; its HNSW index alone used 217 MB. After backup and cleanup, the database settled near 293 MB, a 59.3% reduction.

A reversible half-precision trial improved non-empty Romanized/Hinglish retrieval from 0/14 to 12/14 in that specific probe set. But exact search timed out on most 1,536-dimensional probes, weak sources surfaced, and the complete answer-quality runs scored between 115/120 and 119/120.

Production stayed on title matching plus ranked lexical retrieval. The simpler system was smaller, more predictable, and better aligned with the measured quality gate.

The July hardening cycle moved the curated regression bank from 112/120 to 119/120 and then 120/120 across the tested routes. That bank measures known contracts—tradition labels, prohibited claims, source evidence, unsupported exact citations, internal leakage, and usefulness. It does not prove universal doctrinal correctness. A later model migration passed repository checks and controlled production smokes; we have not found evidence of a fresh full 120-call sweep after that migration.

Architecture decision

Measure new complexity before keeping it

Storage and model spend were optimized independently, without removing the safety layers around the model.

Database size

Before cleanup≈718 MB
After cleanup≈292–293 MB

Unused vector structures accounted for about 217 MB of the removed footprint.

Model cost per answer

Earlier model≈1.2¢
Measured replacement≈0.026¢

The model changed. Retrieval, policy, verification, and human review stayed in place.

07. We accepted visible trade-offs

Safety added latency

In one controlled August 24 canary, complete answers returned in 9.41 seconds on Android, 9.46 seconds on iOS, and 14.48 seconds on web. Those are single-run operational measurements, not a benchmark.

We accepted the slower first response rather than reopen unverified streaming. The next performance work is to reduce evidence packets, parallelize safe retrieval, cache deterministic results, and expand direct coverage without weakening the final gate.

Source grounding is not sentence-level attribution

The product can filter and show allowed source cards, but a supplied source is not the same as a proven citation. We describe this honestly and are building a tighter evidence contract.

The knowledge and application planes fail differently

JainKosh’s MediaWiki remains on Bluehost shared hosting; the AI backend and admin tools run on Vercel, with operational data in Supabase. A healthy model does not guarantee a healthy website.

We mitigated a Bluehost human-check and path-cookie failure, added browser telemetry, outside-in probes, and archived logs, and continue to treat the host’s edge behavior as residual risk, not a permanently solved incident.

08. What changed for readers

The architecture only matters if it produces a calmer product.

Readers now get:

  • shorter, structured answers in their selected language;
  • compact JainKosh source references instead of plausible-looking arbitrary links;
  • explicit fallbacks when exact support is missing;
  • follow-up prompts that continue the study path;
  • feedback that is tied to the answer being reviewed;
  • consistent policy across web and native apps;
  • no visible RAG, routing, QA, or model-planning language.

The deeper change is recoverability. A wrong answer can still happen. It is now easier to identify which layer failed, route it to a human, and make the correction durable.

09. Five product lessons we would reuse

Start with the customer risk, not the technology

“Add RAG” is a solution statement. “Readers cannot distinguish a supported answer from a fluent guess” is a product problem. The second framing exposes better options.

Turn a vague quality problem into failure classes

One hallucination metric could not tell us whether to improve retrieval, source policy, display rules, route parity, or feedback. A failure taxonomy made ownership and testing possible.

Make the goal narrower than the aspiration

Universal accuracy is not measurable. Supported answers, blocked known failures, allowed citations, abstention behavior, and repair time are.

Treat interface states as part of the trust system

Formatting, uncertainty language, feedback confirmation, and source presentation determine whether users can understand and challenge an answer. They are not polish applied after the model work.

Delete complexity that does not win

Vector infrastructure was technically interesting. It did not beat the complete product system on the evidence we had, so we removed it.

What we are doing next

The next bets tighten proof rather than add spectacle:

  1. show which passages the final answer actually relied on;
  2. build a labeled set of difficult questions with expected canonical sources;
  3. measure Hit@1, Hit@3, unsupported claims, latency, and fallback behavior;
  4. calibrate abstention for source-critical questions;
  5. reduce the 9–14 second response range without exposing unverified text;
  6. keep strengthening access, retention, and least-privilege controls around the volunteer workflow.

JainKosh did not become trustworthy because one model got smarter. It became more inspectable, more bounded, and more repairable because product, engineering, and community review were designed as one system.

See the product

Ask a question in Hindi, Hinglish, or English.

Explore the live experience—and keep the same critical eye for sources, clarity, and uncertainty.

Try JainKosh Ask