Cluesora Corpus

Live · free to start

Your organization already knows.
It just can't tell you.

The answer is in a doc someone wrote in March, a thread that scrolled away, a runbook two people ever read. Nothing is missing. It's just unreachable — by your team, and by every agent you've pointed at it.

Corpus turns all of it into something an agent can actually reason over — and keeps it inside the boundaries your organization already has.

3 seats, 200 documents, 2 agent connections. No card, no clock.

WHAT FEEDS IT Docs & PDFs Confluence Notion Web pages AI AGENTS · VIA MCP ChatGPT Claude Claude Code Cursor …and any MCP client your team adopts next add a document, carry on — re-indexed in the background, eventually consistent CLUESORA CORPUS every document, three indexes · every table, queryable · one boundary ENTITLEMENTS CONCEPT TREE leaf → branch → trunk · the same idea in two documents lands in one place ENTITY GRAPH the people, systems, clients and processes that link across documents FULL TEXT every entity searchable on the exact words somebody used DATASETS rows stay rows — pages cite them live, queries total them and show the formula MCP MCP SERVER search · fetch · explore · query · write — scoped to what the person may see SEARCH & ASK answers from your own sources — never the open internet THE PLATFORM ON TOP Cluesora, told in three dialects EDUCATION schools · colleges · coaching ENTERPRISE onboarding · competencies · hiring INDIVIDUAL your own materials, your own pace every session, assessment and score written back into the same Corpus

One Cluesora box, one boundary — everything your organization has written, indexed three ways and held behind entitlements. The learning platform beside it reads and writes the same graph. Hover or tap any piece to follow its flow.

The knowledge didn't leave. The people who could find it did.

Every organization runs on things that were only ever explained once — why the pipeline is configured that way, what the client actually meant, which of the four documents is the one that's still true. It gets written down often enough. It just gets written down somewhere.

So search returns the document instead of the answer. A new hire spends three weeks assembling what a colleague could have said in a sentence. And when you finally hand the whole pile to an agent, it does what the search box did, faster — because a folder of files is not a thing you can reason over. It's a thing you can grep.

What Corpus does

One document. Three indexes. Nobody files anything.

Most systems index a document one way and live with what that costs them. Corpus lays every document down three ways at once — and, more to the point, keeps all three as the document changes. The structure is not a snapshot taken at upload. It is maintained.

ANY DOC CONCEPT TREE leaf → branch → trunk ENTITY GRAPH relations across documents FULL TEXT every entity, searchable ENTITLEMENTS YOUR AGENT sees only its own re-indexed in the background, continuously

Every document is laid down three ways at once — as concepts in a tree, as entities and relations that reach across documents, and as full text over every entity. An agent's question is answered from all three, and never reaches anything the person behind it isn't entitled to see.

Concept tree

Ideas resolved into leaves, branches and a trunk — so two documents that say the same thing in different words land in the same place instead of sitting beside each other forever. The relations between concepts have a direction: what must make sense before this does.

Entity graph

The people, systems, clients and processes your documents keep mentioning, and the relations between them — the connections that only exist across documents, which no single file can hold.

Full text

Every entity searchable on its literal words, because sometimes the right answer is the exact phrase somebody used and no amount of structure beats finding it.

All three are maintained continuously and eventually consistent — the work happens after the write, not during your question. You add a document and carry on. Corpus catches up.

Datasets

A page stores what you concluded. A dataset stores what you measured.

Prose is the wrong place for a price list, a set of readings, a spec sheet — anything whose rows repeat the same fields. Pasted into a page it can only be read again. Held as a dataset it can be filtered, joined, totalled, and checked against a rule.

So in Corpus, rows stay rows. A page cites the rows it needs with one small block, and the rows render live for every reader — the page never holds a private copy that can drift. Ask for a derived number and it arrives with the formula that produced it, rather than as an estimate.

Your agents learn the same habit: search for the table that already holds this kind of row, and extend it. Never a second copy, never one table per note.

PAGE How we priced Corpus ```dataset dataset: corpus-plan-tiers compute: held / per_month ``` RENDERED LIVE Starter 250 / mo 2,000 held Pro 1,000 / mo 8,000 held Max 5,000 / mo 40,000 held the page never holds its own copy — it can't go stale the query rows, live DATASET corpus-plan-tiers one row is one plan tier TIER PER_MONTH HELD Starter 250 2,000 Pro 1,000 8,000 Max 5,000 40,000 COMPUTED · WITH ITS FORMULA held / per_month → 8.0 · 8.0 · 8.0 every paid tier fills in eight months — a query found that

A page keeps the reasoning. The dataset keeps the arithmetic. Neither has a private copy of the other — the page cites its rows with one block, the rows render live for every reader, and a derived number arrives with the formula that produced it.

We priced Corpus this way

6

revisions of the pricing model in one sitting — a price, a cap, a tier, a cap again. Each one was an update to a table and the same three queries re-run, not six rewrites of a document.

8.0

months for every paid tier to fill its workspace — held ÷ per_month, identical on every row. The caps had been written across three revisions in three tables. Nobody designed that. A query found it.

1

cohort double-counted, caught because a total came back at 1,000 users for 500 people. In prose that is a sentence three paragraphs above a table, and nothing collides. Aggregates collide.

Agents · via MCP

Inside Claude, ChatGPT, Cursor and Claude Code — reading before it writes.

Corpus is an MCP server. Connect the agent your team already works in and it can search the concept tree, walk the entity graph, open a page in full, query a dataset with the arithmetic in the query — and write back what it learned, where it belongs.

Every write starts with a search. The agent finds the page that already covers a subject and extends it; finds the dataset that already holds this kind of row and adds to it. The same decision does not get filed twice under two names — which is the failure mode every shared drive eventually dies of.

Reading is never metered. Search, fetch, explore, the map, every query — free on every tier, including Free.

  1. 1 Create a workspace. Free, three seats, no card.
  2. 2 Add what you have. Upload documents, point it at a Confluence space or a Notion page, paste a URL. Corpus reads it in the background.
  3. 3 Connect an agent. Add the Cluesora MCP server to Claude, ChatGPT, Cursor or Claude Code. It signs in as you and sees what you see.

cluesora corpus · the toolset an agent sees

# read — never metered

searchfetchexploreget_knowledge_mapquery_dataset

# write — search first, then extend

save_knowledgecapture_team_factdefine_datasetwrite_dataset

# a query, with the arithmetic in it

query_dataset

dataset: "corpus-plan-tiers"

compute: [{ expression: "held / per_month" }]

→ 8.0 · 8.0 · 8.0 — with the formula attached

Every tool is scoped to the person behind the connection. An agent cannot read a page its user couldn't open.

Entitlements

A memory an organization can actually switch on.

The reason most organizations never point an agent at everything they know is not that the technology can't hold it. It's that nobody can promise what comes back out, to whom.

Corpus assembles every answer inside the asker's boundary. Not a filter bolted on at the end — the entitlement is part of how the answer is built, which is the only version of this that survives the awkward cases: the half-shared document, the contractor with access to one project, the person who moved teams last week and should stop seeing the old one.

An agent inherits the permissions of the person behind it. Nothing else was ever safe to ship.

On Free, the boundary is the workspace itself — everyone is an admin and every folder is shared, which is the right shape for one person or a team trying it. Roles, permission groups and folder-level access begin at Starter.

Built in, not bolted on

  • Multi-tenant from the first commit — not a single-user tool that grew an organization mode.
  • Boundaries that hold on the ambiguous cases, not just the obvious ones.
  • The same rules whether the question arrives from a person, an agent, or an integration.
  • Datasets too — a column above your clearance is named, never silently dropped, so you know what you weren't shown.

There are other memory products. These are the differences that are structural.

Structure it derives, not structure you author

Commonly

Link-based tools ask you to draw the connections yourself, then go stale the week you stop.

In Corpus

Corpus reads what you already wrote and works the structure out — which ideas are the same idea in different words, which sit under which, which reach across documents. And it keeps working it out as the documents change.

Relations with a direction

Commonly

Most memory stores keep flat facts and retrieve whatever looks similar to the question.

In Corpus

Corpus keeps directed relations between concepts — what has to make sense before this does. A similarity search cannot answer that question; it has nowhere to put the arrow.

Numbers that stay numbers

Commonly

A table pasted into a page is a picture of data. The next revision of the page is a second, slightly different picture.

In Corpus

Rows that repeat the same fields live in a dataset. Pages cite them and render them live; queries total them and show the formula. There is one copy, and it is the one you can ask questions of.

Boundaries built for an organization

Commonly

The category grew up around one developer and one vault, so permissions arrived late and thin.

In Corpus

Corpus was built inside a multi-tenant product from the first commit. Entitlements are not a filter applied at the end — they are part of how the answer is assembled.

Not a prototype

This layer is already carrying a product.

Corpus isn't a new idea announced early. It's the substrate underneath everything Cluesora already does — the layer that reads an organization's own materials, works out which ideas depend on which, and follows a single concept from the page it was written on to the person who never quite got it.

That product needed a knowledge base that could hold ordered, derived, permissioned structure at organizational scale. We had to build one. Corpus is that layer, opened as its own door — and when a team on Corpus wants the platform on top, it's an edition change on the same organization. Same graph. No migration.

Pricing

A flat price per seat. Reading is free on every tier.

Every tier carries the whole layer — the three indexes, datasets, the MCP server, Search & Ask. Free is flat: everyone an admin, every folder shared. Boundaries begin at Starter — users, roles, permission groups, folder-level access. From Pro, SSO and the Mr. Cluesora teammate come along. Starter runs to 100 seats; Pro and Max have no limit. Prices in Indian rupees, per seat per month, before GST.

One person, or a team trying it

Free

₹0 forever
  • Up to 3 seats — everyone is an admin
  • Every folder shared with the whole workspace
  • 200 documents and 200 datasets held, shared across the workspace
  • 25 writes per seat a month

Small teams

Starter

₹400 /seat · month
  • Up to 100 seats
  • User management & roles
  • Permission groups & folder-level access
  • 2,000 documents and 2,000 datasets held, per seat
Most popular

Teams that write every day

Pro

₹1,000 /seat · month
  • Unlimited seats
  • Everything in Starter — users, roles, permission groups
  • OIDC SSO & SCIM provisioning
  • Mr. Cluesora in Slack & Teams

Teams running agent pipelines

Max

₹4,000 /seat · month
  • Unlimited seats
  • Everything in Pro — SSO, SCIM, the teammate
  • Query workbench & analytics dashboards
  • 40,000 documents and 40,000 datasets held, per seat
Every allowance, side by side

Starter to 100 seats, Pro and Max unlimited. A data-residency requirement, or invoice billing? Talk to us.

When accountability is grounded in evidence, it stops feeling like surveillance and starts feeling like care.

Why an organization can switch this on

ClueSora Corpus

Point your agents at what you already know.

Create a workspace, add a few documents, connect the agent you already use. It starts reading before you've finished the coffee. Free for three seats, for as long as you like.