The first JainKosh chatbot had a product problem disguised as a model problem.
It could answer in Hindi, Hinglish, and English. It streamed quickly. It sounded confident. But confidence was exactly what made a wrong verse, a mixed tradition, or an unrelated source link dangerous: the answer looked finished before it had earned trust.
For a community knowledge product, that is not a minor quality bug. One polished mistake can make readers question every answer that follows.
So we stopped asking, “Which model hallucinates less?” We asked a more useful product question:
How might we make an unsupported answer difficult to produce, easy to detect, and useful to correct?
That reframing changed the roadmap. The work was no longer a model upgrade. It became a product system: source boundaries, risk-based routing, final verification, visible feedback, volunteer review, and regression tests.
01. The product problem: confidence outran evidence
The early architecture was intentionally small: a JavaScript widget on JainKosh.org called a Next.js API, which sent a language prompt to a general model and streamed the result.
That MVP proved something important: people could ask natural questions without learning MediaWiki navigation. It also exposed three user risks.
- A plausible answer could cite the wrong evidence. One response reused the wrong illustration for a Moolachar verse; another blurred two sections of Chhahdhala.
- A general answer could cross a doctrinal boundary. Digambar-default responses sometimes introduced other-tradition details without the user asking for a comparison.
- A reported mistake could return. Prompt edits improved wording, but they did not create a durable invariant across paraphrases, models, or web and mobile routes.
The prose was polished. The evidence chain was not.
Calling all three failures “hallucination” hid the product decisions we needed to make. Retrieval failures, tradition-policy failures, unsupported citations, generation drift, route inconsistency, internal-text leakage, and broken feedback each belonged to a different layer.
02. The goal: design for recoverable trust
“Zero hallucination” was never a credible product promise. Our goal became narrower and measurable:
- Ground source-critical answers in JainKosh material.
- Block known high-risk errors before they reach a reader.
- Show uncertainty instead of inventing precision.
- Keep the experience consistent across web, iOS, and Android.
- Turn confirmed mistakes into tests so the same failure is harder to repeat.
This is the difference between an output and an outcome. Shipping retrieval, a review queue, or a new model is an output. The outcome is a reader getting an answer that is more supportable—and a volunteer being able to repair the system when it is not.
Architecture decision
The model became one component in a larger evidence and review path.
Early MVP
Fluent output, no enforced evidence boundary
Current path
03. Four product decisions changed the system
Decision 1: give the model a source boundary
We built an offline ingestion pipeline around JainKosh’s MediaWiki pages and uploaded study material. Each source keeps its title, URL, revision identity, checksum, timestamp, and section path. Searchable chunks preserve language and quality signals.
At this review point, the production corpus contains 1,277 active sources and 27,733 chunks.
The live retrieval path is deliberately practical:
- prefer a strong title match when a user names a granth or JainKosh page;
- run ranked lexical retrieval across normalized Hindi, Hinglish, Romanized, and Jain terms;
- remove weak, off-topic, navigation, test, and sample pages;
- send the model a compact packet of text, title, section, URL, and quality metadata.
Retrieval supplies evidence. It does not prove that the draft used every supplied passage correctly.
Decision 2: route by risk, not by one universal prompt
Some questions can use open generation. Others should not.
JainKosh now resolves every question into a typed answer policy. That policy decides whether to use a direct template, guarded generation, source-critical generation, or open generation; whether retrieval and a source line are required; and whether a tradition label or explicit comparison is needed.
Known doctrinal landmines have canonical guards: aliases, prohibited claims, required context, verified source references, and language-specific fallback behavior. A small set of stable, high-risk facts can bypass model improvisation entirely.
An OKF-style concept layer helps recognize aliases, related ideas, and “do not confuse with” boundaries. It guides routing; it is not treated as proof or exposed as a citation.
Decision 3: verify before display
The model’s draft is not automatically the user’s answer.
Before display, JainKosh can reject a prohibited claim, replace an unsafe response with a canonical fallback, prevent unsolicited tradition mixing, remove internal routing text, strip unapproved URLs, flag unsupported exact citations, or say that a source still needs verification.
After internal planning text once reached the interface, we also stopped exposing low-risk answers token by token. The provider can stream internally, but the product buffers the draft, applies the final policy, and only then shows the answer.
Decision 4: keep one answer path across products
Web, iOS, and Android once risked becoming three different AI products. We moved accuracy logic into one shared answer planner and native chat handler. The apps remain native, while policy, retrieval, verification, citations, and answer capture are centralized.
That means one verified correction can improve every surface without waiting for three independent implementations.
Architecture decision
Different failures are stopped at different layers; no single layer claims perfect accuracy.
04. The interface became part of correctness
A trustworthy answer is more than its factual paragraph. Readers also judge the language, formatting, source reference, waiting state, follow-up prompts, feedback controls, and whether the same question behaves consistently next time.
The chatbot lives on JainKosh.org itself—not behind a separate login or app download. Open the widget, pick a language, and ask:
- language modes are explicit, with English, Hindi, and Hinglish selectable directly in the widget;
- long explanations use headings, bullets, and bold terms;
- source-critical answers can name a granth reference;
- follow-up suggestions are phrased as useful next questions;
- quality monitoring has a direct on/off control;
- the same architecture powers desktop and mobile with consistent answer policy.

Product evidence · Mobile experience
The same product works on mobile with the same English answer quality
Mobile view of the same chatbot on JainKosh.org. The question, answer, language selector, monitoring notice, and input are all present. Capture from production on August 24, 2026.
The screenshot is product evidence, not a perfection claim. The mobile view shows the same English answer quality as desktop, but the current system does not yet prove at sentence level which retrieved passage supported every phrase. That gap is on the roadmap.
The same answer policy reaches real phones
The product is not confined to a website. A native iOS app ships to testers through TestFlight, and an Android app ships through Google Play Internal Testing—both using the identical retrieval, verification, citation, and review infrastructure.
The native apps are not wrapped web views. They use SwiftUI and native Android components, with onboarding, a Today daily-practice tab, Jain Calendar with color-coded tithis, full-text search, and an AI tab that calls the same shared answer planner as the web widget. Vikas Ji, Arpit Bhaiya, Shilpy, and several family testers have been testing builds since July.

Product evidence · Native iOS app
The same product also ships as a polished native iPhone app via TestFlight
Native iOS onboarding with English and Deep selected. The same retrieval, verification, and answer policy that runs on JainKosh.org powers the native app. QA capture from iPhone 17 Pro simulator; debug-only footer excluded.
The screenshot shows the English onboarding flow. Every answer that follows—regardless of platform—travels through the same retrieval, policy routing, verification, and feedback capture. One verified correction can improve iOS, Android, and web without waiting for three independent changes.
05. Human review closed the loop
Feedback that disappears into an inbox does not improve a knowledge product.
JainKosh captures the question, final answer, platform, language, model context, source cards, and user response when quality monitoring is enabled. QA and internal activity are marked separately so test traffic does not masquerade as real usage.
A volunteer review queue prioritizes reported answers, negative feedback, incorrect-answer signals, and repeated questions. Reviewers can mark an answer verified, incorrect, needing discussion, or spam; add a category and note; and leave an audit record. “Answer looks correct” is an intentional resolution—review must be able to clear a false alarm as well as confirm one.
At this review point, the system holds 2,062 production answer records and 227 recorded review actions. These are queue records, not user counts.
The final step is the most important: a confirmed correction becomes a regression fixture. The fix is not complete when a prompt changes. It is complete when the relevant router, verifier, citation, parity, or UI test would fail if the bug returned.
Architecture decision
Automation helps diagnose and test. A reviewer still decides what should change.
Evaluation sends new failures back into the next review cycle.
06. We measured complexity—then removed it
A PM’s job is not to approve the most advanced architecture. It is to keep complexity only when it improves the user outcome.
| Choice | What the evidence showed | Product decision |
|---|---|---|
| A bigger model as the primary fix | Model changes did not create a source boundary or prevent known repeat errors | Rejected as the strategy |
| Vector retrieval in production | Better Romanized recall in one probe, but higher storage, timeouts, and weaker full-answer scores | Removed after a reversible trial |
| Unverified token streaming | Faster first text, but unsafe or internal text could appear before checks | Buffer, verify, then display |
| Human review without regression tests | A reviewer could correct one answer while the same class of error returned | Convert confirmed errors into contracts |
The vector experiment is the clearest example. The first implementation pushed database storage to 718 MB; its HNSW index alone used 217 MB. After backup and cleanup, the database settled near 293 MB, a 59.3% reduction.
A reversible half-precision trial improved non-empty Romanized/Hinglish retrieval from 0/14 to 12/14 in that specific probe set. But exact search timed out on most 1,536-dimensional probes, weak sources surfaced, and the complete answer-quality runs scored between 115/120 and 119/120.
Production stayed on title matching plus ranked lexical retrieval. The simpler system was smaller, more predictable, and better aligned with the measured quality gate.
The July hardening cycle moved the curated regression bank from 112/120 to 119/120 and then 120/120 across the tested routes. That bank measures known contracts—tradition labels, prohibited claims, source evidence, unsupported exact citations, internal leakage, and usefulness. It does not prove universal doctrinal correctness. A later model migration passed repository checks and controlled production smokes; we have not found evidence of a fresh full 120-call sweep after that migration.
Architecture decision
Storage and model spend were optimized independently, without removing the safety layers around the model.
Database size
Unused vector structures accounted for about 217 MB of the removed footprint.
Model cost per answer
The model changed. Retrieval, policy, verification, and human review stayed in place.
07. We accepted visible trade-offs
Safety added latency
In one controlled August 24 canary, complete answers returned in 9.41 seconds on Android, 9.46 seconds on iOS, and 14.48 seconds on web. Those are single-run operational measurements, not a benchmark.
We accepted the slower first response rather than reopen unverified streaming. The next performance work is to reduce evidence packets, parallelize safe retrieval, cache deterministic results, and expand direct coverage without weakening the final gate.
Source grounding is not sentence-level attribution
The product can filter and show allowed source cards, but a supplied source is not the same as a proven citation. We describe this honestly and are building a tighter evidence contract.
The knowledge and application planes fail differently
JainKosh’s MediaWiki remains on Bluehost shared hosting; the AI backend and admin tools run on Vercel, with operational data in Supabase. A healthy model does not guarantee a healthy website.
We mitigated a Bluehost human-check and path-cookie failure, added browser telemetry, outside-in probes, and archived logs, and continue to treat the host’s edge behavior as residual risk, not a permanently solved incident.
08. What changed for readers
The architecture only matters if it produces a calmer product.
Readers now get:
- shorter, structured answers in their selected language;
- compact JainKosh source references instead of plausible-looking arbitrary links;
- explicit fallbacks when exact support is missing;
- follow-up prompts that continue the study path;
- feedback that is tied to the answer being reviewed;
- consistent policy across web and native apps;
- no visible RAG, routing, QA, or model-planning language.
The deeper change is recoverability. A wrong answer can still happen. It is now easier to identify which layer failed, route it to a human, and make the correction durable.
09. Five product lessons we would reuse
Start with the customer risk, not the technology
“Add RAG” is a solution statement. “Readers cannot distinguish a supported answer from a fluent guess” is a product problem. The second framing exposes better options.
Turn a vague quality problem into failure classes
One hallucination metric could not tell us whether to improve retrieval, source policy, display rules, route parity, or feedback. A failure taxonomy made ownership and testing possible.
Make the goal narrower than the aspiration
Universal accuracy is not measurable. Supported answers, blocked known failures, allowed citations, abstention behavior, and repair time are.
Treat interface states as part of the trust system
Formatting, uncertainty language, feedback confirmation, and source presentation determine whether users can understand and challenge an answer. They are not polish applied after the model work.
Delete complexity that does not win
Vector infrastructure was technically interesting. It did not beat the complete product system on the evidence we had, so we removed it.
What we are doing next
The next bets tighten proof rather than add spectacle:
- show which passages the final answer actually relied on;
- build a labeled set of difficult questions with expected canonical sources;
- measure Hit@1, Hit@3, unsupported claims, latency, and fallback behavior;
- calibrate abstention for source-critical questions;
- reduce the 9–14 second response range without exposing unverified text;
- keep strengthening access, retention, and least-privilege controls around the volunteer workflow.
JainKosh did not become trustworthy because one model got smarter. It became more inspectable, more bounded, and more repairable because product, engineering, and community review were designed as one system.
See the product
Ask a question in Hindi, Hinglish, or English.
Explore the live experience—and keep the same critical eye for sources, clarity, and uncertainty.
Try JainKosh Ask