Features
25,000 Pages. About 10 Minutes. Why We Built CaseMark Source Differently.
A 25,000-page medical chronology in about 10 minutes doesn't come from a faster model alone. Andraž Krašček explains the architecture behind CaseMark Source: ingestion that makes records usable, an agent layer that delivers a finished report, and open-weight models matched to each job.

A medical chronology should help you understand a case. Getting one should not become a project of its own.
But when the source material is 25,000 pages of medical records, the preparation can become the bottleneck. Before anyone can evaluate the treatment history, someone has to make the documents readable, find the relevant information, organize it, and connect it back to the record.
We built CaseMark Source around a simple question: what would it take to make that work fast and economical, even when the record set is enormous?
Not just a faster model. Not a bigger chat window. A platform designed for the entire job.
Consider a 25,000-page medical chronology produced in about 10 minutes, at a fraction of the cost you would expect from a traditional provider. The interesting part is not just the turnaround. It is the architecture that makes that kind of outcome possible.
The work starts before the AI writes anything
Most conversations about AI start with the model. For document-heavy legal work, that skips a critical step.
The model needs usable information.
Medical records arrive as scanned PDFs, mixed document collections, and files with inconsistent formatting. Giving an AI a pile of documents is not the same thing as giving it a foundation it can reliably work from.
Our intake and ingestion engine builds that foundation. OCR turns scanned pages into machine-readable text. The ingestion pipeline organizes that text into manageable sections and retains page references where available, so the information remains connected to its source.
We then create embeddings: numerical representations that help the system find passages by meaning, not just exact words.
That matters when the same condition or treatment is described differently across records. It also means subsequent work can use an existing searchable foundation instead of starting from an unprocessed stack of PDFs every time.
This is not the glamorous part of AI. It is one of the most important.
Linc turns models into a working system
A language model can generate text. Completing a complex document workflow takes more than that.
The system needs to know which records belong to the task, what kind of report to produce, how to access supporting material, and where to deliver the finished result.
That is where Linc comes in.
Linc is our open-source legal AI agent, built on the Pi agent framework and connected to case.dev's models and legal tools. In Source, it provides the agent layer that works with the selected source materials and workflow instructions to produce a deliverable.
Think of the difference between asking a model a question and giving a capable colleague an assignment. The assignment comes with materials, tools, an objective, and a definition of done.
Source supplies that structure. Linc works within it.
The completion boundary matters, too. A model saying "done" is not the same as a report being delivered. Source's workflow integration requires a successful report handoff, not just a final chat message.
For a customer, that engineering detail translates into something straightforward: the output is the product, not the conversation.
Open-weight models change the economics
We do not believe every step in a legal workflow needs the same model, or the most expensive one.
Reading a document, creating a search index, finding a passage, and synthesizing a report are different jobs. Treating them as one undifferentiated AI request makes it harder to optimize either speed or cost.
Open-weight models are an important part of our approach. They give us more control over how models are deployed and served, including running capabilities on infrastructure we control. Our embedding infrastructure, for example, supports a self-hosted open-weight model alongside other model options.
The point is not to pick a side in an open-versus-proprietary debate. It is to choose the right capability for the work.
That flexibility lets us improve the platform as models improve, without making the customer rethink their workflow every time the AI market changes.
Scale the work, not the waiting
A large record set should not require one enormous prompt or one process doing everything in sequence.
Source separates document preparation, indexing, agent execution, and report delivery. Work can be broken into bounded pieces, with concurrency where the stages allow it and additional capacity as demand grows.
That is the practical meaning of scalable architecture: growth does not have to mean rebuilding the product around the next large customer.
It is not literally infinite compute. It is a foundation designed to expand with the workload.
The economics follow the same principle. Prepare source material once. Reuse the resulting text and index. Match models to tasks. Avoid moving and processing the entire record unnecessarily.
Those decisions compound. Speed and affordability become properties of the system, not competing priorities.
The goal is earlier understanding
A fast chronology is useful because it gives a professional an organized starting point sooner.
It does not replace judgment. The important dates, statements, and conclusions still need to be checked against the underlying records. Preserving source connections is part of making that review practical.
For reporting firms, insurance teams, and legal service providers, the opportunity is bigger than a single faster report. It is being able to take on substantial work without treating every substantial record set as an operational exception.
That is what we are building with CaseMark Source.
The customer should not have to think about OCR pipelines, embeddings, model hosting, or agent execution. They should be able to bring us the records and get useful, reviewable work back.
The complexity belongs in the platform. Not on their desk.


