KANONIC is an AI research engine designed to solve the high hallucination rates in academic AI tasks, which the deck claims reach nearly 40% for citations (Slide 3). The core innovation is the integration of SAML (Security Assertion Markup Language) tokens, allowing researchers to use their existing institutional credentials to grant an AI model access to paywalled, non-STEM, and non-Open Access content. Unlike competitors that rely on manual uploads or STEM-heavy databases like Semantic Scholar, KANONIC logs access for publishers and automates deep research queries. The founding team consist…
Key takeaways
- The platform addresses a 40% hallucination rate in AI-generated research references (Slide 3).
- KANONIC differentiates itself by accessing non-STEM and non-Open Access content without requiring manual user uploads (Slide 4).
- The technical solution relies on SAML tokens via OpenAthens or Shibboleth to authenticate the AI model's access to multiple sources (Slide 5).
- The founding team is composed of three members with backgrounds from Oxford and LSE, specializing in AI, philosophy, and algorithm design (Slide 2).
- The product includes a reporting feature that logs source access for publishers to maintain ethical standards (Slide 7).
- Market strategy involves a three-stage rollout: building version 1.0 with a Cosmos Institute grant, launching in select HEIs, and then a public release (Slide 8).
- The deck identifies an £11B total market, with a £1.5B mid-tier and a £10M initial target (Slide 8).
- There is no slide detailing the specific funding amount requested or the planned allocation of capital.
Slide-by-Slide Analysis
Slide 1: Title Slide
The deck opens with a minimalist title slide: 'INTRODUCING KANONIC: THE AI RESEARCH ENGINE.' It includes a contact email for Elian McCarron. The branding is clean, using a muted palette and a geometric logo. This slide establishes the product category immediately without unnecessary fluff.
Slide 2: Founding Team
The team slide features three founders with strong academic pedigrees. Puyu Wang is an Oxford Engineering PhD researching Semantic Web and LLM applications. Elian McCarron , who holds degrees from Oxford and LSE, leads Operations and Fundraising. Oliver Ogden is another Oxford Engineering PhD focused on algorithm design. The team is positioned as highly technical and academically rooted, which aligns with the product's focus on scholarly research.
Slide 3: The Problem
This slide identifies two core issues: AI hallucination and paywalls. It cites a 2023 study (Athaluri et al.) stating that 40% of references from research questions were hallucinated by LLMs. It also highlights a disparity in Open Access (OA) content: while STEM is near 70% OA, Arts and Humanities (A&H) are less than 25% OA . The slide uses a graph from the 'Hugging Face Hallucination Leaderboard' to visualize the decline in hallucination rates over time, though they remain significant for research tasks.
Slide 4: Existing Solutions
KANONIC uses a competitor matrix to compare itself against Scite, Scispace, ConnectedPapers, Elicit, and ChatGPT. The primary differentiator is the ability to access non-OA non-STEM content without manual user uploads. The slide criticizes current tools for 'STEM dependency' (relying on Semantic Scholar) and 'unethical manual upload' processes that keep publishers in the dark about content usage.
Slide 5: Our Solution
The solution is presented as a three-stage process. Stage 1 involves SAML token authentication via OpenAthens or Shibboleth. Stage 2 allows the AI to access multiple sources using that token. Stage 3 is an iterative search process. The output is a 'detailed research report' evaluated by FINER criteria (feasible, interesting, novel, ethical, and relevant). A key claim here is that KANONIC logs every source access for publishers, creating an ethical trail.
Slide 6: What is SAML?
Recognizing that investors might not be familiar with the underlying technology, this slide defines SAML (Security Assertion Markup Language). It explains that SAML allows for Single Sign-On (SSO) and that academic institutions provide these tokens to students and faculty. This slide justifies the technical feasibility of the product by pointing to the 'huge existing infrastructure' between institutions and publishers like JSTOR.
Slide 7: Features
This slide shows a product mockup of the KANONIC dashboard. Features include AI Web Search integrated with SAML, and a core 'Logging, Reporting & Expiry' system. The 'Model Expiry' and 'Access Logging' features are specifically marketed as 'solved for publishers,' suggesting a B2B2B strategy where keeping publishers happy is as important as serving the researcher.
Slide 8: Market Size & Strategy
The market is visualized with three circles: £11B, £1.5B, and £10M . While the circles are labeled, the specific definitions (TAM/SAM/SOM) are not explicitly typed on the slide, though a link is provided for the data. The strategy is a standard three-step rollout: Build 1.0 (funded by a Cosmos Institute grant), launch in select HEIs for beta testing, and then a public release with a subscription-based model targeting both B2C and B2B (HEI licenses).
Slide 9: Our Mission
The final slide focuses on the company's philosophy. It emphasizes 'Adapting to AI in Higher Education,' 'Aligned by Design' (working with publishers rather than against them), and 'Using AI to Serve, Not Replace.' It explicitly states that the tool is not meant to write reports for users but to enhance 'researcher agency' by reducing time spent on 'rabbit holes.'
What KANONIC Does Well
The deck is exceptionally clear about the technical 'how.' By focusing on SAML tokens, KANONIC provides a credible answer to the question of how they will bypass paywalls without violating copyright or requiring users to pirate PDFs. This 'ethical' angle is a strong differentiator in a market currently dominated by tools that often ignore publisher rights.
The problem definition is also strong. By citing specific hallucination rates (40%) and the lack of Open Access in the Humanities (25%), the founders demonstrate a deep understanding of their niche. They aren't just building another 'AI for research'; they are building 'AI for the Humanities,' where the data problem is most acute.
What is Missing from the KANONIC Deck
The most glaring omission is a specific Ask slide . There is no mention of how much money the company is looking to raise, the valuation they are seeking, or what the milestones will be for the next 18 months. While they mention a grant, investors need to know the capital requirements for the 'Public Release' phase mentioned on Slide 8.
Additionally, the Market Size slide (Slide 8) is under-explained. While it provides a link for data, a pitch deck should stand on its own. The £11B figure is large, but without knowing if that represents the global academic publishing market, the AI software market, or the HEI budget, it feels like a placeholder. Finally, there is no Traction slide showing beta sign-ups, waitlist numbers, or specific university partnerships beyond the 'select HEIs' mentioned in the strategy.
Founder Takeaways: What to Copy
The Competitor Matrix (Slide 4): Instead of just listing names, KANONIC lists specific technical limitations (e.g., 'STEM dependency') that their product overcomes. This makes the 'Why Us' argument much more persuasive. · The Educational Slide (Slide 6): If your product relies on a specific protocol or technology (like SAML) that isn't common knowledge, including a 'What is [X]?' slide is essential to ensure the investor follows your logic. · The Ethical Positioning: In the current AI climate, showing how your tool benefits the data providers (publishers) as much as the users is a smart way to de-risk the investment from a legal and partnership perspective.
Frequently asked questions
- What is the primary problem KANONIC is solving?
- According to Slide 3, the primary problem is AI hallucination in research. LLMs currently perform poorly on citation tasks, with one study showing 40% of references are hallucinated. This is exacerbated by the fact that less than 25% of arts and humanities papers are open access, meaning AI models lack the retrievable source data needed for accuracy.
- How does KANONIC access paywalled academic papers?
- Slide 5 and Slide 6 explain that KANONIC uses SAML (Security Assertion Markup Language) tokens. Users authenticate via institutional providers like OpenAthens or Shibboleth. The AI model then uses this token to access non-Open Access content across multiple publisher databases, such as Oxford University Press or JSTOR, on the user's behalf.
- Who are the founders of KANONIC?
- Slide 2 lists three founders: Puyu Wang (Oxford Engineering PhD, AI Society Treasurer), Elian McCarron (Oxford Philosophy MSt and LSE BSc), and Oliver Ogden (Oxford Engineering PhD). Their expertise spans semantic web applications, fundraising, operations, and energy-efficient algorithm design.
- What is the business model for KANONIC?
- Slide 8 outlines a tiered-pricing, subscription-based model. The company plans to sell individual B2C subscriptions to researchers and group licenses directly to Higher Education Institutions (HEIs). The goal is to make the tool free at the point of use for institution-affiliated researchers through these B2B licenses.
- Is there any evidence of current traction in the deck?
- Traction is limited to the mention of a grant from the Cosmos Institute on Slide 8. The founders state they are 'mid-way building our first version' and plan to conduct comparative analysis against competitors. There are no mentions of active users, revenue, or signed LOIs from universities yet.
