How the main legal AI products charge, and how legal engineers can bring the cost down.
Like almost all the articles I write, this one began with confusion. I kept reading about the cost of legal AI, the use of tokens and legal engineers, and it had all become one mush in my head. So I researched it, and I have written this article both for myself and for you, to bring a bit more clarity to how legal AI is priced and why it matters.
Legal AI has two price tags, and understanding the difference between them is crucial.
A. Seat pricing: This charges you for the number of people who have access to the software. Think of a restaurant. Seat pricing is like a buffet: you pay one price and eat what you want.
B. Token pricing: This charges you for the text the AI processes, both what you send it and what it writes back. In the restaurant analogy, tokens are like paying for each dish. If you order lots of sides and mains, the bill will be higher, but if you only want a side dish, it will be cheaper.
It is crucial to understand that not every legal AI product is priced the same way. Some charge per seat, some charge per token, and others charge a mix of both, so the price a firm pays depends on how the product it chooses is billed.
The table below shows how each of the main legal AI products charges.
| Product | Charges by | How it works |
|---|---|---|
| Harvey | Seat | A fixed price per seat, which Harvey has said it plans to keep.1 Behind the scenes, Harvey buys access to AI models from providers such as OpenAI.2 |
| Astra for Law | Tokens | Offered to legal tech companies, including Harvey and Legora, through OpenAI's API, which bills per token. OpenAI has not disclosed a price for Astra for Law itself.2 |
| Legora | Hybrid | A fixed price for the core platform, plus consumption billing for its newer Agent Pro product: you pay for the work the agent delivers.3 |
| Gemini Enterprise for Legal | Hybrid | A seat that includes a usage quota, with pay per use beyond it. This describes the general Gemini Enterprise, as Google has not published pricing for the legal edition.4 |
| Claude for Legal | Hybrid | The plugins are free, and you pay through your Claude plan. On Team and seat based Enterprise plans, a seat includes a usage allowance, with tokens billed at API rates beyond it.5 |
Lawyers often use AI differently from most people. A lawyer will come in with hundreds of pages of documents and ask the AI to analyse every nuance of them. That level of analysis uses a huge number of tokens, and it can drain a firm's usage very quickly.
Now picture how this often goes wrong. A lawyer uploads 60 documents and types a one line prompt: "do this", then "now do that". The prompts are weak, so the answers need correcting, and the conversation drags on for an hour. In many tools, every message sends the whole conversation, documents included, back through the AI again, and the bill keeps growing.6 For a firm on a token based system, where it pays for every token used, the cost can climb very quickly, all from what should have been a simple review.
The fix is batching. Instead of feeding the documents into one long chat, each document is run as its own separate task, all at the same time. Legal AI products already work this way. In Harvey's Review Tables, each row is a document and each column is a question, so every document is reviewed individually, across hundreds or even thousands of documents.7 Claude's Cowork can run independent tasks in parallel.8 For firms that pay by the token, Anthropic also offers batch pricing through its API: the AI works through the tasks in its own time, usually within an hour, and the firm pays half the normal price.9 Because every document stands alone, the AI is not re-reading a growing conversation each time, which is exactly what made the lawyer's hour long chat so expensive. The trade off is speed, so batching suits big jobs like due diligence and not a lawyer who needs an answer in five minutes.
A legal engineer sits between lawyers and technology. Law firms define the role differently, but the idea is the same: they combine legal knowledge with technology.10 They do not need to code the AI itself. They need to understand how it can be used most efficiently, so that the firm becomes more efficient too.
One area a legal engineer can focus on is advising lawyers on how to use AI in a way that keeps token usage down without hurting the legal work. The aim is to reduce cost and time whilst keeping the quality high.
If legal engineers teach lawyers to use AI well, the result follows naturally. Lawyers who prompt properly and batch their work where it is available, instead of running long chats, use fewer tokens, and fewer tokens means a lower bill. The firms that get the most from AI are likely to be the ones whose lawyers are open to learning from legal engineers.
Legal AI is not priced in just one way. Some products, like Harvey, charge a fixed price per seat, others, like Astra for Law, charge per token, and a growing number, including Legora, Gemini Enterprise and Claude, charge a mix of both. Where tokens are involved, how the AI is used matters: long chats with weak prompts cost more, while separate tasks and batching cost less. That is the gap legal engineers can help to close, by teaching lawyers to use AI well. Understanding how a product is billed is the first step to understanding what it really costs.