50+ AI and machine learning engineers over 1,000+ days, alongside 50+ chartered accountants, 20+ top-tier IP lawyers and 10+ former Income Tax Commissioners. Built on India's public tax and IP record — never on customer documents. The people who spent careers answering assessments taught it what an assessing officer is actually asking.
Anyone can point a language model at a corpus of judgments. What separates a usable engine from a plausible one is two things held together — a machine learning team building the architecture, and qualified practitioners labelling it: this argument worked, that one did not, and here is why the officer asked.
1,000+ days building the architecture — the domain pretraining, the retrieval and grounding layer, the similarity engines and the deterministic reconciliation logic. Not a wrapper assembled in a quarter.
Powers every model in the ensembleOver 1,000 days building TaxEye — annotating notices, mapping grounds to sections, marking which submissions were accepted and which drew a second query, and encoding the reconciliation logic behind every figure.
Powers TaxEye · notices, assessment, appeal1,200 days building Entermark — grading similarity the way a registrar grades it, labelling which Section 9 and 11 arguments carried, and codifying when opposition beats rectification.
Powers Entermark · search, prosecution, oppositionThey sat on the other side of the desk. Their contribution is the part no corpus contains: what an income tax question actually means during assessment and appeal — the concern behind the wording, and what answer closes it.
Powers intent classification across every tax workflowA 142(1) asking for "details of sundry creditors" is rarely about the list. It is a bogus-purchase enquiry, or a Section 68 credit-worthiness test, or a year-end cut-off check — and the reply that closes it differs completely in each case. Former Commissioners labelled that distinction across thousands of real questions, which is why the engine drafts to the concern rather than to the sentence.
Every ITAT order, High Court judgment, examination report and gazette notification is published by the State — hundreds of thousands of documents in which a professional argued a position and an authority accepted or rejected it. That is what the engine learned from, not your client’s bank statement.
A model trained on that learns what actually persuades an assessing officer or a registrar. It does not need to read your client's bank statement to learn it — and so we never had to ask.
RegEye's corpus starts in 2000 — the point at which the Indian government began publishing and amending law at its current pace — and has not had a gap since. Every circular, notification, master direction, press release and authority order, from every regulator we cover.
Completeness is the point. A partial archive tells you what changed; a complete one tells you what a provision looked like in the assessment year under dispute, and which amendment moved it.
Most legal AI asks you to hand over your files so the model can get better. Ours got better because a hundred and thirty engineers and practitioners spent four thousand days building and teaching it on documents the government already published — which is why "we never train on your data" costs us nothing to promise.
Learning happens in separate, isolated layers. The base engine improves on public corpora. Your workspace adapts to your practice without any of it flowing back. Nothing crosses between customers, ever.
Written for the engineer or CTO on the evaluation call. If a term here is doing work in our stack, it is named; if it is not, it is absent.
Different problems want different models. A phonetic similarity judgment and a limitation calculation should not run through the same architecture, and in our stack they do not.
50+ AI and ML engineers over 1,000+ days built the architecture. 50+ chartered accountants and tax practitioners spent 1,000+ days on TaxEye, 20+ IP lawyers from top-tier firms 1,200 days on Entermark, and 10+ former Income Tax Commissioners taught it what an assessment or appellate question is really asking. The engine has also been in continuous training since inception, so there is no knowledge cut-off between the public record and what it recognises.
To the year 2000, without a gap — the point at which the Indian government began amending law at its current pace. That completeness is deliberate: an assessment for an earlier year turns on what the provision said then, so the engine reconstructs point-in-time law and follows the amendment chain forward.
Learning is split into isolated layers. The base engine learns only from public corpora. Your workspace adapts to your firm privately and never contributes to the shared model. Expert corrections improve the base engine as labelled professional judgments, not as copies of client documents. There is no pooling across customers.
Correct — and it costs us nothing to promise, because the training corpus is the public record. The engine learned Indian tax and IP reasoning from half a million published judgments and 75 lakh trademark records. It does not need your client files to do that, and it does not get them.
No. It is an ensemble: layout-aware document understanding, legal reasoning models trained on Indian judgments, phonetic and visual similarity engines, a deterministic reconciliation engine, and a constrained drafting layer. General-purpose models are used where they are genuinely the right tool, inside that structure and never as the whole of it.
Applicability classification is measured against the decision of the professional who verified it. Every circular is reviewed by a lawyer, chartered accountant or company secretary, so each classification carries a ground-truth label. The figure is the agreement rate over that review set.
Indian customer workloads run in India — AWS Mumbai for production, AWS Hyderabad for encrypted backups. Documents are processed in-region and are not sent outside it. Sub-processors, including model providers, all operate within India and are listed in the register available on request.
Yes. Every classification, draft and figure links to its source, and every action, version and approval is timestamped in an immutable audit log exportable in full for regulatory review.
Thirty minutes, your file, our bench — and a security questionnaire answered in writing if your reviewers need it.