Private LLM for law firms and tax advisers on-premise: confidentiality, document volume and GPUs
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A law firm or tax practice with 200 to 2,000 staff runs a private LLM on GPU servers it controls, so client documents, prompts and answers stay under its own access rules, with search over the DMS per matter, long-context contract review, drafting and translation
- At EU level the GDPR and the AI Act apply to the platform, and the CCBE (2 October 2025) and CFE Tax Advisers Europe (25 February 2026) have published guidance on AI use by lawyers and tax advisers
- Document volume sizes the retrieval side: by our arithmetic 10 million chunks of 512 tokens take about 19 hours to embed on one L40S and about 56 hours on one L4 at NVIDIA’s published NIM rates at concurrency 1, and about 41 GB of 1,024-dimension vectors
- Context length sizes GPU memory: with gpt-oss-120b one 32K conversation needs 1.125 GiB of 16-bit KV cache and one 128K contract review 4.5 GiB, so one RTX PRO 6000 holds 19 or 4 and one H200 NVL 55 or 13 by our estimate
- With the example values in this article, a firm of 800 staff needs two servers, each with four RTX PRO 6000 Server Edition or two H200 NVL plus one L40S, so that either server carries the peak alone
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Private LLM for law firms and tax advisers: the short answer
An on-premise LLM for a law firm or tax practice with 200 to 2,000 staff runs on GPU servers the firm controls, so that client documents, prompts and answers stay on hardware under the firm’s own access rules. The platform answers questions over the document management system (DMS) with access per matter, reviews long contracts in one context window and helps with drafting and translation. Hardware follows the document volume, which sizes the embedding and index side, and the context length of reviews, which sizes the GPU memory per conversation. Head count sets only how many conversations run at the peak.
By our estimate, a firm of 800 staff with the example values below needs two servers, each with four RTX PRO 6000 Server Edition or two H200 NVL for the language model and one L40S for the retrieval models, so that either server can carry the peak alone. The figures are memory estimates from published model configurations, not measurements, and a pilot replaces the usage values with the firm’s own.
Confidentiality, GDPR and the AI Act at EU level
Professional-secrecy obligations for lawyers and tax advisers are set nationally, and how they apply to an AI platform is for the firm’s own legal assessment. At EU level, the European bodies of both professions have published guidance. The Council of Bars and Law Societies of Europe (CCBE) published its guide on the use of generative AI by lawyers on 2 October 2025. eucrim’s report of 31 October 2025 summarises its confidentiality point as “Lawyers may not input personal, confidential, or client-related information into GenAI tools.” The report does not say whether this point also covers tools that run on hardware under the firm’s control.
CFE Tax Advisers Europe published its Charter of Tax Advisers’ Rights and Obligations in an AI-Influenced Tax Advisory Environment on 25 February 2026. Under data protection and confidentiality it states: “Compliance with GDPR must be respected, as well as professional secrecy rules and contractual duties when using external tools.” Under autonomy and professional scepticism it adds that “Tax advisers retain full responsibility to accept, adapt or reject AI outputs.”
The GDPR applies to the client and staff data in the platform as to any other system. Article 5(1)(f) requires processing “in a manner that ensures appropriate security of the personal data”. Article 32(1) names measures such as encryption, the ongoing confidentiality of processing systems and the ability to restore access to the data after an incident, and Article 28 governs any provider that processes the data on the firm’s behalf. Files in employment, family or personal injury matters can hold data concerning health, a special category under Article 9(1).
Under the AI Act, a firm that uses an AI system under its authority is a deployer (Article 3(4)). Article 4 has applied since 2 February 2025 and, as amended by Regulation (EU) 2026/1744, asks deployers to take measures to support the development of AI literacy of their staff. Annex III, point 8(a), lists as high-risk the systems intended to be used by a judicial authority or on its behalf to assist it “in researching and interpreting facts and the law and in applying the law to a concrete set of facts”, or to be used in a similar way in alternative dispute resolution. Regulation (EU) 2026/1744 did not amend Annex III, and Article 113 as amended applies the high-risk rules to Annex III uses from 2 December 2027. Whether a particular use in a firm falls under point 8(a) is part of the same legal assessment.
Law firm workloads and the GPU each one needs
The language model in our examples is gpt-oss-120b, released under Apache 2.0 with 117B parameters of which 5.1B are active; its model card places it in use cases that “fit into a single 80GB GPU”. Its config.json gives 36 layers, half of them over a 128-token sliding window, with 8 key/value heads of dimension 64. The 16-bit KV cache therefore costs 36 KiB per token in the full-attention layers: 1.125 GiB for a 32K conversation and 4.5 GiB at 128K, the model’s maximum of 131,072 tokens. With the rule from our private ChatGPT server sizing guide, one RTX PRO 6000 leaves about 22.2 GiB for cache beside the 60.8 GiB of weights, and one H200 NVL about 62.5 GiB.
| WORKLOAD | MODEL CLASS | MEMORY PER UNIT | GPU IN THE EXAMPLES |
|---|---|---|---|
| Search and Q&A over the DMS | 1B embedding model and reranker | small weights; throughput decides | L40S or L4, or a 24 GB MIG instance |
| Drafting and summaries | gpt-oss-120b, 32K context | 1.125 GiB per conversation | RTX PRO 6000: 19; H200 NVL: 55 |
| Contract review at 128K | gpt-oss-120b, full context | 4.5 GiB per conversation | RTX PRO 6000: 4; H200 NVL: 13 |
| Document translation | 9B to 12B translation LLM | 16-bit weights on one card | L40S or RTX PRO 5000 for about 1 million words a day |
| Indexing the archive | embedding model in batches | throughput per card | L40S: 144.3 chunks of 512 tokens per second |
Our estimates for gpt-oss-120b with a 16-bit KV cache, one copy per card: 0.9 × driver-visible memory, less 3 GiB, less 60.8 GiB of weights; layers and heads from the model’s config.json on Hugging Face, read on 10 October 2026. Embedding rate from NVIDIA’s NIM 1.14.0 performance page (updated 30 September 2026, Llama Nemotron Embed 1B, FP8, batch 64, concurrency 1); translation from our machine translation guide.
With another model the figures change and the method stays the same. Our embedding and reranker guide compares the models and serving engines. For translation, our on-premise machine translation guide sizes servers by words per day and lists which translation models carry non-commercial licences.
Document volume: sizing the index for the DMS
The retrieval side is sized from the DMS archive of pleadings, contracts, correspondence, tax returns and working papers. Take, as example values, 1 million documents split into 10 chunks of 512 tokens each, which gives 10 million chunks. NVIDIA’s NIM 1.14.0 performance page lists 144.3 texts of 512 tokens per second on an L40S and 49.2 on an L4 for its 1B embedding model in FP8, at batch size 64 and concurrency 1. By our arithmetic the first full indexing run takes about 19 hours on one L40S and about 56 hours on one L4.
The model returns vectors of 384 to 2,048 dimensions. Set to 1,024 dimensions at four bytes per value, each vector takes 4 KiB, so 10 million chunks need about 41 GB before index structures, chunk text and metadata. Vector databases that store quantised vectors need less, and the index sits in server RAM unless the database keeps it on disk, which sets the memory of the retrieval server. Every chunk also carries the matter number, the client, the document date and the access list of its source, which the filter per matter needs.
After the first run only new and changed documents are embedded, a small daily load. A change of the embedding model means embedding the whole archive again, so the index card is sized for that run as well. Scanned files need text recognition before they can be embedded, which adds its own processing time.
Long contracts: context length decides the memory
Reviewing a full agreement with its schedules in one request needs a long context window, while a question across the archive goes through retrieval. At 128K tokens, one review keeps 4.5 GiB of cache on gpt-oss-120b, four times a 32K chat, so three reviews running together take as much memory as twelve chats. Documents longer than the model’s 131,072 tokens are split or handled through retrieval.
A 128K prompt also needs compute before the first token appears, so the response time of reviews is checked on the target card with the firm’s own contracts. When several lawyers ask questions about the same contract, prefix caching reuses the processed prompt. Our long-context LLM hardware guide compares the cache per token of other models and the configurations for 256K and 1M tokens.
Worked sizes for 300, 800 and 2,000 staff
We use the example values of our sizing guide: 40 per cent of staff active in the busiest hour, 6 requests each per hour, 30 seconds per request and a peak factor of 2. For 800 staff that gives 320 users, 1,920 requests an hour and 32 conversations at the peak. On top we assume long reviews at the peak, 2, 6 and 15 for the three sizes, as example values. Each size places two servers that each hold the whole peak, so that the service continues if one fails.
| STAFF | PEAK CHATS AT 32K | REVIEWS AT 128K | CACHE AT PEAK | EACH OF TWO SERVERS |
|---|---|---|---|---|
| 300 | 12 | 2 | 22.5 GiB | 2 × RTX PRO 6000 or 1 × H200 NVL, plus 1 × L4 |
| 800 | 32 | 6 | 63 GiB | 4 × RTX PRO 6000 or 2 × H200 NVL, plus 1 × L40S |
| 2,000 | 80 | 15 | 157.5 GiB | 8 × RTX PRO 6000 with the L40S in a third server, or 3 × H200 NVL plus 1 × L40S |
Users and peaks from the example values above; reviews at the peak are example values; cache at 1.125 GiB per 32K chat and 4.5 GiB per 128K review for gpt-oss-120b in 16 bits, with 22.2 GiB per RTX PRO 6000 and 62.5 GiB per H200 NVL by our estimate.
At 300 staff the peak needs 22.5 GiB of cache, just above one RTX PRO 6000, so each server takes two cards, or one H200 NVL with room to spare. At 800 staff four RTX PRO 6000 hold 88.8 GiB and two H200 NVL 125 GiB against the 63 GiB needed. At 2,000 staff eight RTX PRO 6000 hold 177.6 GiB, which fills an eight-GPU server, so the L40S for the retrieval models goes into the H200 NVL variant or a third, smaller server. Every row is a memory ceiling; response time under load needs its own test.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users, and check the rack, power and airflow before we quote. Send us your head count, document volume and the contract lengths you review through the form below.
Access per matter, logging and retention
In a firm the unit of access is the matter. The assistant may only retrieve passages from documents the requesting user can open in the DMS, and when a lawyer is screened from a matter, the index follows the change. Our guide to RAG on company data explains how the filter works at retrieval and why a prompt instruction does not control access. The table maps the controls a firm asks for to the platform and to the EU texts above.
| CONTROL | IMPLEMENTATION | EU REFERENCE |
|---|---|---|
| Access per matter | each chunk carries the matter and the access list of its DMS source; retrieval filters before ranking | GDPR Art. 5(1)(f), Art. 32(1)(b) |
| Information barriers | a screened user’s rights change in the DMS, and the index re-syncs on that change | GDPR Art. 32(1)(b) |
| Query and answer log | user, time, matter, retrieved passages and answer, kept on the firm’s storage | GDPR Art. 5(2) |
| Retention | prompts, answers and logs erased after the period the firm sets, test copies included | GDPR Art. 5(1)(e), Art. 30(1)(f) |
| No external services | public model APIs off; any external call only by explicit decision, visible in the log | GDPR Art. 28, Art. 44 |
| Restore | backup of index, configuration and logs, with restore tests | GDPR Art. 32(1)(c) and (d) |
| Verification | answers cite the source document and its date for review by the lawyer | CCBE guide; CFE Charter |
Regulation (EU) 2016/679 on eur-lex.europa.eu; CCBE guide as reported by eucrim on 31 October 2025; CFE Charter of 25 February 2026. The mapping is our technical reading, not a legal assessment.
Article 30(1)(f) asks the record of processing to state, “where possible, the envisaged time limits for erasure”, so the retention period of prompts and logs is set before go-live.
Assistants built in our Private AI/ML service cite the source and respect each user’s access rights, with queries and answers logged. Describe your DMS and how matter access is granted in the form below.
What we supply
We supply the RTX PRO 6000 Server Edition, the H200 NVL, the L40S, the L4 and the RTX PRO 5000 discussed in this article, as cards for servers you run or in AI servers built to order with 2 to 8 GPUs per node, burn-in tested, with manufacturer warranty, on one EU contract and invoice. We check the rack, power and airflow before we quote and return a configuration and quote within one business day. The platform on top, with RAG, access rights and the query log, is our Private AI/ML service with engineering by our partner Vixen.UNO, on your infrastructure or on dedicated hardware in a Tier-3 data centre in Lithuania. NVIDIA AI Enterprise licences come on the same invoice, and our professional GPU range lists the cards.
FAQ
What does an on-premise LLM for a law firm need?
Can a law firm use a confidential LLM under professional secrecy?
How many GPUs does a law firm GPU server need for 800 staff?
What hardware do tax advisers need for on-premise AI?
How big is the vector index for a law firm DMS?
Is legal AI on-premise high-risk under the AI Act?
Send us your head count, the number of documents in your DMS, the contract lengths you review in one request and how matter access is granted today. We reply within one business day with a configuration and quote, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day