The evidence gate for enterprise AI.
HalluciProof is designed to check each factual claim against its supporting evidence, apply a risk-based decision before delivery, and create a verifiable record of the result.
Addressing the Enterprise AI Verification Gap
Generative AI produces articulate answers, but lack of granular verification exposes organisations to compliance and remediation risk.
The Granularity Problem
Generative AI can produce fluent answers containing incorrect figures, unsupported conditions, unverifiable citations or statements based on outdated sources.
A single aggregate score for the whole answer may not show which individual statement is unsupported. When validation only happens post-delivery, organisations are left handling costly customer complaints, regulatory scrutiny and remediation.
The Pre-Delivery Control Layer
HalluciProof is developing a dedicated control layer that sits between generative AI systems and the end recipients of their answers.
Rather than inspecting output in retrospect, the platform isolates material factual claims, tests evidence sufficiency against verified records, and enforces organizational policy before answers are delivered.
HalluciProof complements existing AI evaluation and observability tools by introducing claim-level evidence lineage, independent adversarial challenge, and signed decision records.
Kainat Zanab
Kainat Zanab holds an MSc in Digital Business Innovation and Management from Lancaster University, awarded with Distinction. The HalluciProof project draws directly on her postgraduate dissertation research into AI hallucinations and enterprise-safe AI deployment.
Relevant Background & Research:
- An MSc consultancy engagement with SAP on AI governance, compliance automation and regulatory risk.
- An IBM research collaboration delivered as part of her MSc.
- An NHS Trust Lancashire AI in Healthcare university project.
- A Cisco, Lancaster University Management School and NTNU university collaboration.
- Enterprise Technology Consultant experience at BT Group, including participation in a pilot cohort trialling AI-assisted tools.
Kainat leads strategy, product definition, evidence methodology, customer discovery and commercial development, supported by a planned founding AI engineer.
Modular Architecture for Verifiable AI Delivery
A 10-module control system structured across four architectural layers, designed for granular evidence testing and auditability.
Input
Captures draft model responses, retrieved grounding context, metadata, and workflow identifiers from enterprise systems.
Core Verification
Decomposes text into discrete factual claims, builds evidence links, deterministically checks figures, and tests sufficiency.
Decision & Assurance
Applies customer risk thresholds, executes runtime gates, and generates tamper-evident cryptographic receipts.
Integration
Connects with vector stores, agent frameworks, SDKs, webhook sinks, and enterprise review consoles.
Claim Decomposer
Separates material factual statements from tone, caveats and non-factual language. Categorises claims including eligibility, prices, dates, citations and general assertions.
Evidence Graph Builder
Links claims to specific passages, records or tool outputs, retaining precise evidence references, source versions and cryptographic content hashes.
Evidence Sufficiency Engine
Tests whether linked evidence materially supports the claim. Evaluates supported, partially supported, contradicted, missing or stale states with deterministic checks for numbers, dates and identifiers.
Cross-Model Adversarial Verifier
Challenges high-risk claims that passed earlier checks using a model from an independent provider and architecture family, avoiding shared inductive bias with generating and sufficiency models.
Risk Policy Engine
Applies customer-defined policies and thresholds by use case and claim type. Higher-risk statements (such as binding financial commitments or legal rights) face stricter evidence requirements.
Runtime Decision Gate
Combines evidence results, challenge outcomes and customer policy to execute deterministic actions: allow, qualify, escalate or block the claim prior to end-user delivery.
Evidence-Decay Monitor
Tracks source freshness over time and triggers re-verification of dependent claims when a grounding document, database record or guideline becomes stale or changes.
Signed Evidence Receipt Service
Hashes, time-stamps and digitally signs claims, evidence references, challenge outcomes and gate decisions for portable and independent verification.
Challenge and Test Generator
Builds targeted evidence-gap and adversarial test suites for each deployment and continuously feeds the customer workflow benchmark.
API, SDK and Connectors
Provides planned interfaces for retrieval pipelines, agent frameworks and enterprise sources. Planned interfaces include REST and streaming APIs, Python and TypeScript SDKs, webhooks and connectors.
Supporting Capabilities
Enterprise controls and reporting utilities designed to accompany core gate verification:
Planned Development Roadmap
Capabilities are currently planned or in development. Month 1 is modelled as January 2027 following endorsement and formal incorporation.
Core Batch Pipeline
Delivery of core batch processing pipeline and execution of the first evidence-assurance assessment pilot.
Runtime MVP
Deployment of runtime MVP and basic human review console for pilot partner testing.
Adversarial Verifier GA
General availability of Cross-Model Adversarial Verifier add-on for high-risk claims.
Signed Receipts GA
General availability of Signed Evidence Receipt Service with cryptographic hashing.
Decay Monitor GA
General availability of continuous Evidence-Decay Monitor and stale source alerting.
Sector Packs & Scale
Agentic trace coverage, regulated sector packs, and enterprise private deployment options.
Advanced Ecosystem
Opt-in federated benchmark exchange, rotating challenger model pool, and multimodal claim verification.
Eight Steps From Draft to Verifiable Decision
How an answer moves from raw generative output to isolated claims, sufficiency checks, and pre-delivery gate decisions.
Capture the draft
Receive the AI answer, retrieved context, and workflow identifier directly from the generation pipeline.
Separate the claims
Identify and classify the individual material factual statements, separating facts from tone and filler.
Connect the evidence
Link each claim to supporting passages, enterprise records or tool outputs with content hashes.
Test evidence sufficiency
Check support, partial support, contradiction, missing evidence, and freshness with deterministic validation.
Independently challenge high-risk claims
Use a separately sourced model family to challenge claims that passed initial checks and identify ungrounded assumptions.
Apply the customer’s risk policy
Evaluate against configured thresholds to return allow, qualify, escalate, or block.
Record and review
Create decision records and, once released, signed receipts. Route escalations to reviewers and log outcomes.
Monitor evidence freshness
Once the decay monitor is released, recheck dependent claims when their supporting evidence becomes outdated.
Selective Challenge
The independent challenge is selective and targeted at high-risk claims, not mandatory for every low-risk statement.
Direct Block on Missing Evidence
Missing or contradicted evidence can trigger block or escalation immediately without requiring an independent challenge.
Customer Policy & Oversight
The customer defines the risk tolerances, sets policy thresholds, and retains full human oversight at all times.
Pre-Delivery Gate Decisions
Every evaluated claim resolves to one of four deterministic gate outcomes:
Release Claim
Release the supported claim under the configured policy when evidence fully corroborates the statement.
Conditional Release
Release the claim accompanied by necessary conditional phrasing, context, or required qualifying disclaimers.
Human Review
Send the claim, linked evidence, and reason to a human reviewer prior to delivery for manual verification.
Halt Delivery
Stop the claim from being delivered to prevent unverified, contradictory, or high-risk claims from reaching users.
Customer Pilot & Adoption Journey
Adoption is designed to follow a phased, low-friction validation path:
Early pilots assess logged responses in batch to measure baseline verification rates. The initial pilot focuses on assessment and does not promise the immediate deployment of the complete future runtime platform.
Target Markets & Launch Pricing
Proposed Year 1 launch pricing from the business plan. Subscriptions are invoiced annually in advance. All amounts in GBP, excluding VAT.
Financial Services
Customer assistants and adviser tools producing factual statements about products, pricing, cover, and eligibility terms.
Legal & Professional Services
Legal research, document drafting, regulatory citations, and structured contract or clause summarisation.
Healthcare
Bounded knowledge assistants using provider-supplied guidance and clinical protocols. HalluciProof supports provider safety processes; it does not provide clinical advice.
AI Vendors
Embedding assurance gates directly into third-party AI software products via API/SDK and white-label licensing.
Evidence-Assurance Pilot
A focused engagement to measure factual support and baseline gate performance on real logged responses.
- Scope: One bounded workflow over 2 to 4 weeks.
- Dataset: Normally 300–500 human-reviewed benchmark claims.
- Criteria: Success criteria agreed formally before the pilot commences.
- Deliverable: Comprehensive benchmark report covering agreed performance measures.
- Methodology: Early pilots assess logged responses in batch.
Payment Schedule: 50% payable on signature and 50% payable on completion.
- One gated workflow
- Core verification engine
- Review console
- Standard receipts when released
- Email support
- Up to five workflows
- Core verification engine & review console
- Agent trace coverage when released
- Evidence-decay monitor when released
- Single sign-on
- Named success lead
- Unlimited workflows
- Private deployment options according to the roadmap
- Custom policies & thresholds
- Audit exports
- Priority support
Optional Add-ons & Vendor Licensing
Specialised capabilities and partner programme pricing (all prices exclude VAT):
Cross-Model Adversarial Add-on
Billed monthly in arrears. Premium verification for high-risk claim categories using independent model architectures. Planned general availability: Month 12.
Regulated-Sector Assurance Pack
Billed monthly in arrears. Pre-configured policy packs for financial services, legal, or healthcare governance. Planned rollout in Years 2–3.
API/SDK & White-Label Licence
£42,000 per year, invoiced annually in advance. For AI vendors embedding verification into their products. Planned vendor programme from Year 2.
Frequently Asked Questions
Key details on platform capabilities, architecture design, data governance, and pilot engagements.
HalluciProof is developing an evidence gate that sits between an organisation’s generative AI system and the people receiving its answers. The platform decomposes AI responses into individual material factual statements, tests whether each claim is sufficiently supported by verified evidence records, challenges high-risk statements using an independent model family, applies customer-defined risk policies before delivery, and generates verifiable decision records.
The platform is currently at the concept and architecture stage, equivalent to Technology Readiness Level 2 (TRL 2). The capabilities described on this website are planned or in development. Month 1 of commercial development is modelled as January 2027 following endorsement and incorporation. We are currently engaging potential pilot partners to validate requirements and benchmark logged responses.
No. HalluciProof is designed to complement existing AI evaluation frameworks and observability tools. Where observability platforms monitor offline metrics or aggregate whole-answer scoring, HalluciProof focuses specifically on pre-delivery runtime gating: isolating claim-level evidence lineage, performing deterministic sufficiency checks, challenging high-risk claims, and issuing signed audit receipts.
When evidence is missing or contradicts a claim, the Evidence Sufficiency Engine flags this directly. Missing evidence or outright factual contradictions can trigger an immediate BLOCK or ESCALATE decision under the customer’s risk policy, without needing to pass through the secondary cross-model challenge.
Cross-model adversarial verification is an optional check applied selectively to high-risk claims that have passed initial sufficiency tests. It employs a model from an entirely separate provider and architecture family to identify potential reasoning gaps or counter-arguments. This cross-family separation prevents shared blind spots between the generating model and the verifier.
A signed evidence receipt is a cryptographically hashed, time-stamped, and digitally signed record containing the isolated claim, evidence references, challenge results, and gate decision. A signed receipt demonstrates record integrity and proof of policy execution at a specific point in time; it serves as an audit trail rather than a philosophical guarantee of objective factual truth.
The planned platform architecture is designed around data minimisation and UK-region infrastructure, incorporating encryption in transit and at rest with strict access controls. By default, full underlying source documents remain within customer systems, with only relevant passages, text hashes, and claim statements passed to the verification pipeline. (Note: These reflect planned architectural controls for the production platform, not functions of this frontend landing page).
The Evidence-Assurance Pilot is a fixed-scope engagement covering one bounded workflow over 2 to 4 weeks. It normally evaluates a sample of 300–500 human-reviewed benchmark claims against agreed success criteria. Early pilots assess logged responses in batch, producing a detailed benchmark report on claim support, contradiction rates, and policy calibration. Payment is structured as 50% on signature and 50% on completion.
All core subscription tiers (Starter at £21,000/yr, Professional at £51,000/yr, and Enterprise at £114,000/yr) are invoiced annually in advance. Monthly amounts shown on the site represent monthly equivalents for ease of comparison and do not indicate a monthly subscription payment plan.
Yes. Under the Year 1 design-partner offer, the full £15,000 fixed pilot fee can be credited against the customer’s first-year annual subscription if they proceed to sign an annual contract within 60 days of pilot completion.
No. Evidence assessment is not a philosophical guarantee of absolute factual truth, nor does HalluciProof guarantee regulatory compliance. The platform provides a rigorous, verifiable technical control layer that tests statements against available source records. Customers always retain ultimate accountability for their communications, workflows, and regulatory obligations.
No. This website is a frontend-only demonstration. The "Request a Pilot" form stores your entered information strictly inside your browser's localStorage on your local device. It does not send any data to external servers, APIs, databases, or third-party CRM services.