Private Document Processing Requires Controlling the Agent’s File Access
Filip Makraduli of Superlinked argues that sensitive document processing should happen before a coding agent sees the file: open models running on infrastructure the user controls can handle OCR, extraction and redaction, then return a smaller artifact for the agent to use. His Obsidian and MCP workflow is designed to reduce exposure to external model APIs and avoid spending context on raw documents, while letting teams choose or adapt models for specific tasks. But a private processing path does not alone guarantee privacy if the agent can read the original file first.

Move private document work out of the agent’s context
Filip Makraduli started with a practical tension: his files lived in an Obsidian vault as local Markdown, but the useful models for working across those files were rented through external APIs. Letting an agent read a private document could send its contents—and any personal information in it—to a model provider. Sending whole documents also consumes context and tokens on work that could be done before the agent sees the result.
Makraduli’s proposed division of labor is to leave the agent in place for ordinary work, while offloading sensitive or expensive document processing to models running on infrastructure the user controls. A document goes first to a service in the user’s cloud. Open models parse it, extract information, or redact personal data; the agent receives a smaller artifact, such as clean Markdown, to continue working with.
The intended benefits are connected. Processing a scan before the agent sees it can reduce the material sent into the agent’s context. Returning a compact artifact can preserve room in that context for the actual task. And keeping the raw document and processing models within a private environment can avoid sending the document to an external model API.
Makraduli also wants the processing layer to be adaptable. Sensitive, specialized tasks may call for models tuned to a particular domain, rather than one general-purpose model. His design therefore treats privacy, cost, and choice of models as a single infrastructure problem: the system must make it practical to run and change open models, not merely offer a local alternative to an API.
A smaller artifact can carry the useful information
The cost argument is especially clear for documents that contain scans, images, or formatting an agent would otherwise have to handle. Makraduli described the expense of repeatedly sending a long document as both a token problem and a context-management problem: even if the full document is available, filling the context with material that could have been processed separately leaves less room for the agent’s other work.
A slide in Superlinked’s workshop materials estimated that offloading a document and returning Markdown could reduce tokens by roughly 85 percent. The slide reported different measurements for two models, comparing text plus page images with the Markdown artifact. Another slide gave a different Opus figure while retaining its Sonnet figure. These are measurements presented for the demonstrated setup, not a universal reduction for every document or workflow.
The core mechanism is simpler than the figures: do not make the agent repeatedly ingest every page and image if another process can turn the document into text first. Optical character recognition, extraction, and redaction can happen on the private processing side. The agent can then work from the result, rather than from the original file.
Makraduli emphasized that using open models for this work need not mean accepting poor OCR quality. He said open-source OCR models are already very good for the task, so handing OCR off from a frontier model does not necessarily trade away quality. More generally, he sees open models as increasingly useful in production, especially when an operator needs to test several models or fine-tune one for a narrow use case.
The tools in the demonstration cover more than conversion to Markdown. docs_to_markdown turns documents such as PDFs and scans into clean Markdown; it uses document-processing and OCR models. extract_entities identifies items such as people, organizations, dates, and amounts. redact_pii replaces personal information with placeholders. describe_image captions or tags a screenshot or photograph. For more structured work, extract_structured returns information against a schema the user defines, while generate_structured produces constrained text or JSON. The remaining tools, summarize_document and answer_questions, provide a short summary or answer questions about document content.
The point of a common processing layer is that an agent can request an operation—such as “redact this”—without choosing the underlying model itself. The server can select the model and pipeline. Makraduli also described the tool set as open to extension: the MCP package is open source, and contributors can add tools while continuing to use the inference layer.
The privacy boundary is an architectural choice
Filip Makraduli built the integration as an MCP server rather than a one-off plugin. A plugin that runs inside an agent may put the document and cluster credentials in the same environment as the agent. His alternative places an edge service in the user’s cloud, where it holds the credential for the inference cluster, accepts document-processing requests, and returns the resulting artifact. The agent connects to the edge; it does not need the cluster key.
The system has three parts. The MCP client is part of the agent host, such as Claude Code or another MCP-capable application. The SIE MCP edge is the server the client calls, and is intended to run in the user’s cloud. The SIE cluster is the inference layer, with a gateway, work queues, GPU workers, and models. In the demonstrated architecture, the edge authenticates requests and routes document jobs to the cluster. The raw bytes go to the edge and then to the in-cloud processing system; the smaller output comes back to the agent.
MCP supplies a shared interface between agent applications and tools. Makraduli’s case for it is operational as well as technical: one server can expose the same capabilities to different MCP clients, instead of requiring a separate integration for each agent. The interface exposes tools rather than specific models, leaving model choice and execution on the server side.
In this implementation, the server uses tools only. MCP can support other forms of interaction, including asking the host model to generate a response or asking the user for input, but Makraduli said the SIE edge does not use those capabilities. It does not call back to the host or user. The client and server negotiate capabilities when connecting; tools are then discovered through the protocol. For a hosted server, the implementation uses Streamable HTTP, with a bearer token for the connector. The cluster credential remains on the edge.
The separation is configurable. The workshop used a managed cluster for the live demonstration, but Makraduli said users can deploy their own cluster in a private environment. He pointed to the open-source SIE repository and its deployment materials, including Terraform and Kubernetes support, as the route to doing so. A local setup is also possible, though he described it as slower to get running. The materials shown also distinguish the connector secret presented by the agent from the SIE API key held by the edge.
These distinctions matter because “self-hosted” can refer to different parts of a system. The edge, cluster, and GPUs need to be placed and configured in a way that matches the privacy requirement. Makraduli’s architecture is meant to let an operator keep the document-processing path inside their own cloud; the workshop’s managed service was a convenience for trying the demo, not a requirement of the design.
The NDA demo shows both the workflow and its limits
Filip Makraduli began the live example with synthetic files in an inbox, including contracts and invoices. One NDA contains signatories’ names and email addresses, along with a contractor’s phone number, Social Security number, birth date, and home address. Makraduli asked the agent to redact the NDA. The resulting Markdown kept the contract’s structure and replaced personal details with typed placeholders such as [PERSON_3], [EMAIL_1], and [SOCIAL_SECURITY_NUMBER_8].
The processing output reported ten redacted spans: three people, three emails, and one each for phone number, Social Security number, date of birth, and home address. The replacement map was not returned, so the placeholders were not reversible through the result. The agent could then work with a version that preserved the document’s basic content while masking those values.
The demo also showed a scanned services agreement. Its contractor details—including a name, contact information, Social Security number, birth date, and home address—were embedded in an image rather than available as ordinary text. The processing path produced Markdown through OCR and replaced personal details with placeholders, while retaining other information such as the contract value and term. This demonstrates why preprocessing can matter beyond token reduction: a document may need OCR before its contents can be extracted or redacted.
But the demonstration also exposed a practical constraint. In a terminal note, the agent said it had reused a previously parsed Markdown file rather than passing the 112 MB PDF through the MCP tool as base64. Reproducing such a large binary input by hand would be unreliable; a script using the SDK or gateway, or a CLI wrapper, would be a more suitable way to supply the PDF bytes. In other words, the document tools may support scans, but getting a large file reliably into the processing path is itself an integration task.
The redaction output also carried a warning: detection is model-based and best-effort, not a certified scrubber. The agent’s displayed note recommended spot-checking before relying on the result for compliance. Makraduli described the output as successful for the example, but the warning is part of the demonstration’s actual contract with the user: automated detection does not establish that every sensitive span has been caught.
The inference layer does the heavy lifting
Filip Makraduli described the MCP edge as a relatively small front end to a more involved inference system. SIE connects a stateless gateway to queued work and GPU workers. The gateway receives a request and publishes it to a work stream; workers pull from the pool and batch requests before running models. Makraduli said that batching at each worker, with workers drawing from a shared queue, performed better in their testing than routing requests to workers first and letting each worker batch only its local slice.
A slide presented an illustrative comparison of 51 percent GPU efficiency for “route then batch” and 89 percent for “pool then batch,” described as 1.8 times the throughput per GPU at the same latency. These are figures from the workshop slide’s illustrative model, not a general performance guarantee. Makraduli framed the shared-queue approach as one source of end-to-end performance, particularly when different jobs or model states would otherwise leave individual workers with poorly filled batches.
The cluster is also designed to run multiple models across a smaller number of GPUs. Rather than dedicating a GPU to every model, the system can keep frequently used models resident and evict colder ones when capacity is needed. Makraduli said that this is particularly useful for smaller models handling tasks such as named-entity recognition, embeddings, or short generation. It can also make it practical to compare task-specific models or LoRAs without permanently assigning each one its own GPU.
These are infrastructure choices as much as model choices. A workload may use a larger model for some tasks and a smaller one that can run on cheaper hardware for others. Worker pools can be configured for different hardware needs, and autoscaling can reduce capacity when it is not in use. The model catalog reflects another operational detail: model families may require different runtimes because their architectures do not all implement the forward pass in the same way. The catalog also distinguishes model roles such as encoding and scoring, the latter including reranking.
Makraduli’s broader point is that better inference infrastructure can improve latency and throughput even if the model itself has not been optimized. For production systems, he argued, the way work is queued, batched, and scheduled can matter as much as the model selected for an individual request. The same inference layer can support document parsing, embeddings, reranking, and generation, rather than requiring a separate serving arrangement for each task.
Adaptation and routing are opportunities, not finished features
Filip Makraduli sees model adaptation as a reason to keep the inference layer flexible. A team could try several task-specific LoRAs—a way to adapt a base model—and compare their performance on a specialized workload. He cited colleagues’ experiments with German legal-language LoRAs as an example of relatively inexpensive adaptation that improved results in a niche domain. The presentation showed one experiment with legal retrieval, where a LoRA improved an in-domain score but performed worse on general German tasks; Makraduli’s verbal point was that teams can test these adaptations quickly and decide whether they suit their own use case.
More ambitious model routing remains on the roadmap, not part of the current implementation. The proposed direction is to choose the model for a request based on its task, expected quality, cost, and potentially which models are already loaded on a GPU. A tool such as document parsing would no longer necessarily map to one fixed model: the system could select a cheaper model for easy jobs, or a more capable one where the task calls for it.
Makraduli also described the possibility of escalating a difficult request, or using multiple models to draft and assess answers. In the workshop slides, these appear as future patterns, not working features in the demo. Such approaches add inference work and introduce their own failure modes: a router can choose poorly; a panel of models consumes additional GPU capacity; and a judge can be biased or less capable than the models it evaluates.
Evaluation is therefore a requirement, not a finishing step. The slides distinguish embedding evaluations from generation evaluations: an embedding benchmark does not answer which generator a router should select or whether a multi-model “council” improved the answer. Makraduli cautioned against spending GPU capacity on models that reason without producing useful output, relying on biased judges, or optimizing a metric that models can game. The infrastructure makes experimentation possible, but it does not make the results trustworthy by itself.
A successful tool call does not settle the privacy question
Filip Makraduli faced the sharpest question in the workshop when an attendee challenged the system’s central assurance. If an agent is already running in Claude Code and has access to the files, could it read the original document before deciding to call the redaction tool? If so, routing the tool call through a private MCP edge would not by itself guarantee that the raw file stayed out of the agent’s context.
Makraduli initially answered that the agent would receive only the redacted version. The questioner pressed on the step before that: the agent must decide to call the MCP tool, and might first read the file as part of making that decision. Makraduli acknowledged that this behavior needs to be tested and that additional guardrails could reinforce the boundary. He said his own tests had appeared to work without leakage when the MCP and the relevant processing were hosted in the private cloud, but he did not present that as a guarantee. As a stronger alternative for a concerned operator, he suggested building a custom harness with an open-source model and implementing the tool calling directly.
That distinction is important: keeping the inference cluster and MCP edge private protects the processing path, but does not automatically prove that an agent with filesystem access cannot inspect a raw file before using the tool. The workshop did not demonstrate a mechanism that removes this possibility from every agent setup. The privacy boundary depends on how access is configured and how the agent behaves, not just on where the GPU workers run.
A second question clarified deployment ownership. The managed cluster in the demo was hosted by Superlinked, but Makraduli said the cluster can instead be deployed in the user’s own private environment using the repository’s infrastructure setup. The MCP server sends the document to the worker for processing; it is not described as a file watcher that independently scans the user’s filesystem.
For a sensitive workflow, the practical questions are therefore specific: Where are the raw files accessible? Can the agent read them directly, or only submit them to the processing tool? Where do the MCP edge and inference cluster run? Which credentials are available to the agent, and which stay on the edge? How will redaction quality be checked? The workshop’s architecture offers ways to place and separate components, but operators still need to verify that those choices enforce the boundary they intend.
