SYCTRA

找一条条款要二十分钟。现在只要几秒,且附带来源。

您的文档——PDF、Word、Excel、扫描件、演示文稿——转化为可用自然语言查询的知识库。每条答案都标注其出处的文档、页码与段落。

  • 答案标注来源且可核验
  • 按业务语料调优
  • 可扩展至 GraphRAG

The information already exists. It is simply not usable.

In most organisations the document estate is a dormant asset: assembled, paid for, never read again.

01

Scattered

Shared drive, mailboxes, three cloud spaces, archives, local disks. Nobody knows which version governs, and two copies of the same document circulate side by side.

02

Unstructured

Scanned PDFs, tables pasted into slide decks, annexes that amend a contract without saying so. None of it fits into a conventional database.

03

Partly obsolete

Some of the estate is no longer valid, and nothing flags it. Searching without separating current from expired produces confident, wrong answers.

What happens between your question and the answer

The engine does not guess: it searches, compares, checks, then writes from what it found.

01 / 05

Ingestion and normalisation

Documents are read whatever their format, scans included via OCR. Tables, annexes and attachments keep their structure. The date, version and origin of each item are retained — that is what will later separate current from expired.

Two levels of use

Direct querying is the starting point. The agentic layer is what turns search into work done.

Direct querying

You ask, the engine finds the information and answers with its source. “What commitments does this contract contain?” gets an answer in seconds.

Document comparison

Two versions, two contracts, two bids: the agent identifies the real differences, not just wording changes, and lays them out line by line.

Structured extraction

Pulling a usable table out of a document estate: expiry dates, amounts, termination clauses, owners, in the format you impose.

Summary and report

The agent reads several documents, structures the information to your template and produces the note or report, with sources appended.

Deliverable generation

A Word document on your template, a slide deck, a table: the deliverable comes out in the expected format, ready to review, not to reformat.

Multilingual corpora

The engine indexes and queries document estates in several languages at once, and answers in the language of the question.

Tuned per corpus, because a contract does not read like a support ticket

“Tuned” means the algorithm itself changes with the corpus: passage weighting, chunking, version handling, date treatment. It is not a cosmetic setting.

Legal / regulatory profile
Priority to exact wording, article numbers and the version in force at a given date.
Catalogue profile
References, variants and technical characteristics, where one digit changes the product.
Corporate document profile
Procedures, minutes and internal notes, where the most recent version governs.
Temporal profile
Corpora where chronology carries the meaning: decision histories, chains of amendments, incident series.
GraphRAG
Graph modelling when relationships between entities matter as much as document content.
Deployment
SaaS hosted in the European Union, on your cloud, or entirely on your own infrastructure.

Three document estates, three tunings

01 · 背景

Twelve thousand engagement documents scattered between the server and team folders. Finding a precedent takes twenty-five minutes, when you know it exists.

02 · 我们实施的内容

Document layer with the corporate profile, isolation by team, and structured extraction to assemble reference files.

03 · 结果

The precedent is found in under two minutes, with the original deliverable and the engagement name. The estate becomes an asset again.

What we measure at pilot

The error rate is measured on your documents, with your real questions, before any industrialisation.

search time on the questions covered by the indexed corpus
-90 %

search time on the questions covered by the indexed corpus

of answers citing document, page and paragraph
100 %

of answers citing document, page and paragraph

between pilot launch and the decision to industrialise
6-10 weeks

between pilot launch and the decision to industrialise

这些量级来自我们在可比范围内于试点阶段的实测。在贵司语料上,它们会在规模化之前实测,而不是事先承诺。

现在就能看到运行的产品

这些产品已经构建完成。部分可在线访问,其余可按需演示。

在线演示

Ultra RAG

读得懂你的文档、并会标注出处的检索引擎。

检索增强搭配知识图谱,按主题精调:法律、企业文档、时间线。每个答案都能追溯到原始段落。

在线试用
在线演示

Legal-Index

把法律检索交给自主智能体。

法条与判例检索、合同分析、起草辅助。每个答案都指向具体条款,以及你关心的那个日期上的现行版本。

在线试用

常见问题

Native and scanned PDFs, Word, Excel, PowerPoint, plain text, images containing text, exports from business tools. Scans go through an OCR layer before indexing.

It links entities together — a contract and its amendments, a company and its subsidiaries, a standard and its versions. Without it you find documents; with it you understand what changes their scope.

Date, version and origin are captured at ingestion, and the search profiles use them. On a legal corpus, the version in force at a given date wins. That assumes the estate can be dated — otherwise, that is the first project.

Never. The corpus stays isolated by organisation, team and user, and the agent's rights are those of the person asking.

Yes, on your infrastructure or private cloud, including with self-hosted models when corpus sensitivity requires it. That is the standard case on regulated estates.

From a few hundred to several hundred thousand documents. Beyond a certain volume the question is no longer size but the quality of the estate and the granularity of access rights.

本方案不做什么

We do not deploy a document layer on a false or obsolete estate. An excellent engine on an expired corpus produces wrong answers that are sourced and convincing — worse than no engine at all. We say so at scoping and start with the cleanup, which is a project in itself. Nor do we claim one hundred per cent reliability: we promise traceable answers and a measured error rate.

How many documents are asleep on your server?

Bring a representative sample to the diagnostic. We will tell you what is usable as is, and what needs cleaning first.