TL;DR: OpenAI released GPT-6 Astra on September 3. It is the company’s new model for complex, multi-step work. That means reading across a file, following a template, checking figures, filling a form, and moving through a desktop job. For law firms, I think the largest gain is at the paralegal’s desk. Start with document processing that produces a reviewable packet with visible gaps. Test one workflow with planted errors, then across completed matters before live use. Measure the time saved after review.
A client has sent bank statements with several months missing. A draft agreement needs checking against the latest asset schedule. Someone still has to turn a collection of emails into a chronology the attorney can use.
That is where I’d start a law firm’s Astra pilot.
The distinction is drafting versus processing. Drafting produces new language. Processing moves existing information through a job until someone can use it. Astra’s gain is in that second pile, and that pile sits on the paralegal’s desk.
I’d train the paralegals before the associates.
Here are ten tests drawn from documented capabilities. I haven’t run all ten. Several workflows existed before Astra.
The document work
Astra is the model. ChatGPT Work handles longer tasks and finished deliverables. Codex handles technical work. Available tools and permissions depend on the account and workspace.
1. Create editable documents using the firm’s templates
ChatGPT Work can create and edit documents using supplied templates and reference material. OpenAI describes Astra as its strongest model for adhering to existing templates.
Start with the approved template and a verified matter-information sheet. Give it an instruction like this:
Preserve the approved boilerplate. Populate matter-specific fields only from the supplied sheet. Mark unresolved fields. Save a separate Word file and provide a list of missing information.
Then compare the output against the master template. Check that the standard language stayed intact and that the changes came from the approved source.
A clean-looking document can hide an altered clause. The comparison is part of the job.
2. Turn a document dump into a usable inventory
ChatGPT can extract information from supported files and analyze material across documents.
Use those capabilities to inventory a defined batch. Ask for each document’s date range and description, with a reference back to the file. Flag possible duplicates without removing anything.
In a family-law matter, the result might show which monthly statements arrived for each account. The paralegal can check the apparent gaps and prepare a targeted client request.
Require an accounting of every input file. Scanned pages and embedded images may need different processing, depending on the tools and plan.
A summary that quietly drops three unreadable documents is an unfinished inventory.
3. Build chronologies that lead back to the evidence
For the chronology test, require a source reference for every material entry. Distinguish the event date from the date someone described it.
An email written on June 12 might describe a conversation on June 8. A useful chronology preserves both dates where relevant.
When two records disagree, keep both accounts with their supporting passages. The assignment is to expose the conflict, leaving its resolution to the reviewer.
Ask for a sortable spreadsheet alongside a short narrative summary. The spreadsheet supports investigation. The summary helps the attorney get oriented.
Every material entry should be easy to trace and check.
The matter behind the documents
4. Reconcile financial information across a document set
In a case study published by OpenAI, Legora reported that its Astra-powered agent reviewed 41 documents in a single financial-statement reconciliation run within minutes. It found all four planted errors, including a £500,000 gap in a revenue note, and recorded its checks.
This was Legora’s configured agent using Astra through the API. It wasn’t a demonstration of an ordinary ChatGPT conversation.
Legora reported nearly 40% improvement on that workflow. Across its broader legal benchmark, the average improvement was about 3%.
That spread is the reason to test each workflow separately.
For a family-law pilot, compare a draft asset schedule with the supporting statements. Require a record of each comparison and a list of unresolved differences. The paralegal checks whether the correct records were supplied and whether an apparent discrepancy has a straightforward explanation.
5. Check whether the documents agree with each other
ChatGPT supports comparing documents and applying criteria from one file to another.
Use a proposed settlement agreement and its approved term sheet. Check whether the payment terms match. Compare the listed assets against the supplied schedule and flag references to exhibits missing from the review set.
Require the exact passages on both sides of every reported difference.
“The payment terms appear inconsistent” leaves someone with another search task. Showing the two provisions gives the reviewer something to resolve.
Also require a list of items that couldn’t be checked. Missing support belongs in the report.
6. Build discovery response matrices and deficiency lists
For this test, supply a defined request set and a controlled collection of potentially responsive material. Ask for a matrix connecting each request with candidate documents, then an internal deficiency list.
A request for two years of financial records might match statements covering only eighteen months. An email may appear responsive while its attachment is absent.
The distinction between candidate and approved must stay explicit. The system should organize the first pass. The team determines responsiveness and makes the privilege decisions.
Require unresolved work to remain attached to the relevant request. A polished matrix should make the holes easier to see.
7. Prepare a source-linked attorney briefing packet
For deposition preparation, start with supplied testimony and exhibits. Ask for a witness summary, followed by proposed areas for questioning tied to specific passages. Keep factual observations separate from suggested questions.
For hearing preparation, request the relevant record references and an issue list. Treat legal-authority research as a separate assignment.
The paralegal verifies the references before the attorney develops the strategy.
OpenAI reports fewer hallucinations on its internal benchmark. That is encouraging evidence about the model, not an error rate for your matter files.
Check factual claims against the record. Verify legal authorities against the actual decisions and their current treatment. A plausible citation is still another item to check.
The repeatable work
8. Move verified information into forms and other applications
Astra’s computer-use capabilities include filling online forms and working in applications. Access depends on the environment and permissions.
One test would start with an approved intake-information sheet. In a supported environment, have the system enter those details into the intake application, then present the completed fields for review.
Check the result field by field, including which record received the information.
Preparing a form and submitting it should be separate permissions. The same applies to drafting a message and sending it.
In tests that were largely adversarial, OpenAI found Astra’s written reasoning harder to monitor than the prior model’s. A log of clicks is not a reason.
If you can’t log what an agent did, don’t give it a high-stakes action.
Keep the pilot on outputs you can check against a source.
9. Run preparation work on a schedule or after an event
ChatGPT Work supports event-triggered tasks for eligible accounts, including triggers from new Gmail messages. Scheduled tasks can also handle recurring work. Approval requirements continue to apply.
A narrow pilot could review qualifying messages in an approved account and prepare private draft acknowledgments for a paralegal. Another could assemble a recurring status brief from authorized sources.
Specify what starts the task and what makes it stop for review.
One current limitation matters: OpenAI says a scheduled task created in a Project cannot access uploaded files or files stored in that Project. Test the sources available to the actual scheduled run.
Start with private preparation. Unattended client messages and court submissions require separate approval.
10. Turn a senior paralegal’s process into a reusable workflow
Skills can package reusable instructions and examples, with code where appropriate. Availability and sharing depend on the product and workspace.
A senior paralegal could define the firm’s chronology process, including required source references and rules for conflicting dates. An example would show what an acceptable entry looks like.
The same approach could capture the document-inventory process or the first-pass discovery matrix.
For a small custom utility, Codex supports building and testing software. One application might check filenames against an index or identify missing required fields before a packet moves forward.
The paralegal defines the process and judges the result. Technical help handles the code. That gives the firm a way to preserve practical expertise in a form others can use.
What the time is worth
Suppose five paralegals each recover four hours a week after review and rework. Across 48 working weeks, that creates 960 hours of capacity.
The firm then has to decide what that capacity is for.
Under hourly billing, less time on a task can mean a smaller bill. Recovered time might support more matters or better economics on fixed-fee work. It could also reduce a backlog. None of those outcomes happens simply because the draft arrived faster.
Measure the full job. A draft produced in ten minutes that takes two hours to repair tells us very little about productivity.
Plant four, find four
Legora’s test offers a useful starting method. Planted errors expose specific failures. They cannot tell you about every error you didn’t think to plant.
Here is how I’d start.
Choose one workflow with an experienced paralegal. Define an acceptable output and who reviews it. Limit the pilot to a named group using approved data and tools in a firm-approved business workspace. OpenAI says Business and Enterprise data aren’t used for training by default. Review retention and connected-app terms separately, along with matter-specific restrictions. Approval to use the model shouldn’t amount to permission to connect every system in the firm.
Plant four errors in a test copy of a completed matter. Keep the answer key out of the model’s inputs. Count what it misses and what it wrongly flags. Plant four, find four. Passing earns a broader test, not live access.
Run the workflow across 20 representative completed matters with read-only source files. Compare against checked results. Measure omissions and review time. Use the findings to decide whether a supervised live pilot is justified. Keep approval gates in place.
Twenty matters is a practical starting point, not proof of safety. The review must also test whether the system accounted for everything it was given.
The person who knows what belongs in the packet should help define the test.
Astra is a paralegal tool first. Train for it that way.
If you enjoyed reading this, please share it with others and subscribe (on smithstephen.com) if you’d like copies sent directly to your inbox (it’s 100% free).
In the US it is a holiday weekend, and week 1 of college football. In full transparency, I’m a huge Oregon Ducks fan. Watched the game yesterday with my Dad, and Magnus was right in the middle of the fun. He also got a great new harness which he is loving for his walks!




