The phrase “without uploading the whole folder” describes a data-flow preference, not an absolute security result. A useful design separates source storage, preparation, retrieval, AI transmission and retention. Central Brain keeps the first three stages in a local project workflow by default. Passages that a connected AI tool retrieves cross into that provider's boundary.
This makes Central Brain relevant to a team that wants to prepare the same project material once, search it repeatedly, and decide when selected context should be used by an AI assistant. It may be a poor fit when the required workflow is browser-only, centrally hosted with no desktop component, or governed by controls the organization has not verified.
What problem should the workflow solve first?
Start with a bounded retrieval problem, not with a promise that AI can understand an entire shared drive.
A good pilot might ask for the termination clause in a known set of agreements, the latest approved requirement in a project folder, or the spreadsheet row supporting a reported figure. Those questions have known sources and make failure visible. A vague request to “know everything” hides missing files, stale versions, weak OCR and ambiguous language.
Place only the authorized pilot material in a dedicated source location. Keep originals unchanged, choose a separate prepared-output location, and document which people and processes can access both. Central Brain's file tools are read-only: they read and search permitted project files and do not modify, move or delete the source documents.
How does Central Brain process the selected files?
Central Brain prepares supported files, divides them into searchable chunks, embeds them with a bundled MiniLM model and stores a local index for the default workflow.
Supported formats are PDF, DOCX, TXT, Markdown, CSV and XLSX. Optional OCR can recover text from scans and photos when it is included in the build. Images, audio and video get metadata-only indexing, which is not the same as a semantic interpretation of every media file.
Semantic search can retrieve passages whose meaning resembles a question, while exact text search is useful for identifiers, quoted phrases, names and precise wording. Results include the source path, prepared-file path, page and chunk where available. The original file remains the source to review.
What information leaves the local project boundary?
In the default workflow, preparation and search stay on storage you choose; selected context leaves when a connected AI tool or an optional cloud connection is used.
Central Brain uses Keygen for licensing, and Keygen receives licensing and device data, not documents. If you enable an optional cloud connection, such as cloud embeddings, or keep project folders in synced storage, the data selected for that capability is handled within that additional boundary.
A connected AI assistant also has its own processing and retention behavior. Central Brain supplies retrieved context; it does not control the provider after transmission. Review the canonical security and data-boundary guide and the connected provider's terms before using confidential or regulated material.
What does a realistic private-file pilot look like?
A realistic pilot uses a separate project, a small authorized set of files, known-answer questions and an explicit review of every outbound connection.
First, record the source-folder owner, retention rule, backup location and the people allowed to see the files. Second, prepare the project and check the prepared output for extraction or OCR problems. Third, run exact and semantic questions whose correct sources are already known. Fourth, follow each result back to the file and page. Fifth, send only non-sensitive test context through the intended AI connection and confirm the receiving account and provider settings.
Repeat the test after changing, adding and removing representative files. A retrieval system is useful only when the team understands how the index changes with the files. Keep a short acceptance record: what worked, what failed, which formats need manual review, and which questions are not suitable for automated retrieval.
Which selection criteria matter most?
Choose on data boundaries, inspectability, retrieval quality, source traceability, file support, operational ownership and the actual connected-tool workflow.
Confirm where originals, prepared output, vectors, memory, backups, logs and credentials live. Confirm whether a person can inspect the stored material and trace a result to the original. Confirm how changed and removed files are handled. Confirm the supported operating system and the exact AI tool the team intends to connect.
Finally, decide who owns project setup, access review, backups, deletion, license administration and incident response. Local storage transfers responsibility to the organization; it does not remove responsibility. For the next evaluation scenario, see source-traceable document search.
Validation checklist
What should you verify before relying on the workflow?
- Use a dedicated folder containing only authorized pilot material.
- Test PDF, document, spreadsheet, scan and changed-file cases that match real work.
- Run questions with known answers and verify the original file and page.
- Inventory licensing, AI connections, optional cloud connections, synced storage and backups.
- Document unsupported formats, extraction errors and questions requiring manual review.
Common questions
Can Central Brain search files without uploading the whole folder to a cloud AI provider?
Yes. In the default workflow, preparing, indexing and searching files run locally, and files, prepared output and vectors stay on storage you choose. Connected AI tools still receive the context they retrieve.
Does Central Brain make confidential files automatically safe for AI use?
No. The customer must control devices, folders, accounts, backups, connections, permissions, retention and the suitability of each file for the intended use.
What business-file formats can Central Brain prepare?
Central Brain prepares PDF, DOCX, TXT, Markdown, CSV and XLSX, with metadata-only indexing for images, audio and video; OCR is optional and depends on the build.
What is the safest first test?
Use a small authorized project with known-answer questions, verify every result against the original, and review all outbound connections before expanding the set of files.
Next steps
- Security and data boundaries: what stays local and what a connection sends
- Central Brain product page: the canonical product and feature description