Since I was a legal aid attorney in 2019, I have been lucky enough to be a part of a welcoming academic community focused on AI and law. The two anchors of this field are both international conferences: JURIX and ICAIL. (I have never attended a companion conference, JURISIN, which is centered in Japan and East Asia). Both conferences have space for nontraditional authors such as myself, but the rules to play by are not always obvious for those of us who didn’t get the traditional research training of a PhD program.
This community is focused on computer science researchers, but with plenty of space for legal practitioners, programmers, and topics in access to justice.
As you can imagine, this community has gotten new interest and attention since the launch of ChatGPT made the field much more accessible.
My selfish interest in this crossover is to bring more research into the hands of practicing lawyers, and to have more researchers get into contact with practicing lawyers so they understand the real-world impact of their work.
In case you, too, are a practitioner and you would like to participate in one of these conferences or a workshop associated with the main conference (like our series of AI for Access to Justice workshops), I thought it might be helpful to explain what I have learned about how to succeed in this space.
What are ICAIL and JURIX?
JURIX, The Foundation for Legal Knowledge Systems, is an annual conference and tends to attract more participants from Europe. In contrast, ICAIL, the International Conference on AI and Law, has traditionally been a biannual conference and attracted roughly equal participants from both North America and Europe. Despite the traditionally Western focus: Asia, South America, Africa, and Oceania have strong representation at both conferences. ICAIL has recently changed to an annual event.
Both conferences have been around for a long time—about 40 years. They have seen significant evolutions in the field of AI and law, from a rules and formal logic-centered field to one that has expanded to include machine learning and generative AI. Generative AI, in particular, has grown to a significant proportion of the presented papers, without displacing formal logic.
JURIX is a smaller conference (usually 2 days of paper presentations and one day of workshops). ICAIL is typically 3 full days of paper presentations and 2 days of workshops, including a PhD consortium.
JURIX is always in Europe, although it has been hosted in several countries. The most recent editions of ICAIL have been in Europe and the United States; the next edition will be in Singapore.
What should I expect if I attend?
Papers at both conferences are written and presented in English. There are good opportunities for informal conversation and connections at breaks (usually every hour and a half), lunches, and a group dinner.
Paper presentations are a unique format. You will get a long set of proceedings to read that includes every paper in the conference, but each author will probably get 10-20 minutes to go through the key points of their paper live. Each paper also has a few minutes for questions. Slides are a common presentation format, and often cover the highlights of the paper in a straightforward manner. Questions can bring the paper to life. I find that the questions are typically sincere and add value (and are not sniping or overly critical) even where there is healthy debate in the field.
Compare this to a typical conference for legal practitioners, where the presentation is usually more visual, independent of any written materials, and panels are the most common format. The AI and law conferences are heavier on substance than charisma and presentation style, both for better and for worse.
You’ll be absorbing a massive amount of information in a short time. Some papers on very technical topics might be confusing (even though I took logic classes as an undergraduate, the formal logic papers are the most complex for me). I get value from all of the presentations and always come away with new insights, even if I don’t follow every technical detail.
The presenters are going to be a mix of professors, PhD students, and some practitioners. This might be the first presentation for some of them.
All of the conferences typically have at least one longer invited talk for each day. The invited talk is equivalent to a keynote. This will usually be given by a leader in the field with strong presentation skills.
The workshops (usually either the day before or day after the main conference) can be more of the same format, or include more diverse kinds of participation such as interactive workshops or demos.
What can I write about?
This is the most important question to answer: what do you have to say for this audience? A good submission adds something new to the field.
Your paper has to touch on one of the topics of the conference. Examples that might resonate with practitioners include generative AI, a data-science centered empirical study, rules-based expert systems, and natural language processing. Innovative applications are welcome at both conferences, although a paper without a study to go with it might be harder to pitch.
Consider:
- An analysis or evaluation of a project that you worked on. You might need help with the statistics from an expert. Talk to an expert early in case it affects the data you collect through the project.
- You might find a researcher on a similar topic who likes to participate at ICAIL or JURIX who is happy to collaborate with you. You can look through proceedings from past years to find out!
- A research study with interviews of court users or other participants in the legal system. This can be qualitative or quantitative.
- A framework that has recommendations for fair, ethical, or reliable use of AI.
- A survey, metanalysis, or review of research in the field that helps everyone make sense of what others have been doing.
- Detailed case studies that walk through effective applications of AI and offer clear insights and recommendations for the future.
- Qualitative and quantitative studies are both fine!
Papers for this audience will be more practical and engage with data more than law review articles typically will.
This is just a sampling of ideas. Looking at the call for papers from a recent conference might offer more ideas.
The important thing is to be thoughtful about the data you gather, explain the limitations of your work, and be able to support the conclusions you draw.
What have you written about for this community?
I tend to write about projects that I have built. I like to share, and it gives me a secondary value for the work. I have also participated in more theory-heavy papers.
- Substantive Legal Software Quality: A Gathering Storm? – second author on this paper about making sure guided interviews (expert systems) are accurate.
- Beyond Readability with RateMyPDF: a deep-dive into the usability of court forms with our team at Suffolk LIT Lab, and a tool that offers automated rating to make your court forms better.
- Weaving Pathways for Justice with GPT: a paper describing the Assembly Line Weaver (a tool to help build Docassemble interviews with the LIT Lab team) and showing that large language models can help with drafting guided interviews for Docassemble.
- Getting in the Door: Streamlining Intake in Civil Legal Services with Large Language Models: a paper describing Lemma‘s MOTenantHelp.org and a way to use LLMs to decide if a tenant is eligible for help based on a legal aid program’s current intake priorities, collaboration with Maastricht Law and Tech Lab.
- That’s so FETCH: Fashioning Ensemble Techniques for LLM Classification in Civil Legal Intake and Referral: a solo study with partners of Lemma that shows that combining multiple LLMs can result in highly accurate classification of a user’s legal problem, facilitating intake and referral.
Do not be afraid to bring an outsider’s perspective or to write about what you know
It can be beneficial to have an outsider perspective at these conferences. RateMyPDF got a best paper award and was nontraditional for the conference. My three more recent papers were more straightforward and practical and probably most interesting to people trying to solve the exact same problems as me.
Running a study
Not every paper needs a study, but it can make sense to include one with a case study or project writeup, and might not be as hard as it seems. My colleagues Jaromir Savelka and Hannes Westermann are both great at thinking of ways to add a study to a practical project.
Start with a research question
What do you want to prove or disprove with your study? Start with a strong research question. The answer to your question should be meaningful to the field. It might drive investments into different approaches or change decisions about which kind of projects to avoid. Yes or no questions can be strong candidates: “does this work?”
You could have two or three research questions. More than that is probably not going to fit in a 10 page paper.
Review the literature
Your paper will not be taken seriously if you do not do some reading about the work that others have already done first. You do not need a full literature review, but spend some time to figure out what other researchers have to say about your topic. Try to find 20-40 papers that are relevant to your topic.
Using ChatGPT or Gemini’s Deep Research modes can be a good starting point, as can a Google search! Once you have found a few key papers that cover the topic you are interested in, look through the bibliographies and dig deeper. Research papers often form networks. The key papers in a field will start to appear in multiple reference lists.
I organize my references in Zotero.
Now, you will start to have an idea of the terminology used to talk about your research questions. Do some keyword searches in research databases. Google Scholar is free and will have a lot of the latest papers, which will often be preprints on arXiv.org. It is OK to cite to preprints in this field, especially for cutting edge applications of AI to legal problems, but pay attention and make sure that you include high quality peer reviewed research in your list as well.
Look at the proceedings for past conferences as a key source. Reviewers will already know the most important papers from JURIX and ICAIL. It’s important to engage with the community you want to publish in.
In computer science, conferences are a common publication venue. They are considered as reliable and perhaps more prestigious in many cases than other journals.
There is no substitute for reading the papers you want to cite, but it might be helpful to put your sources into a tool like NotebookLM to help you decide where to start. Later, NotebookLM can help you find relevant support for your claims and identify gaps in your analysis.
Again: it is important that you read the papers you cite and add additional papers to your pool beyond any initial search! AI is a truly useful tool for this step and can be useful as a starting point and to expand your search, but it won’t catch everything, and it often will miss the most significant papers in a field. I have found that running a deep research will get me a lot of papers, about 50% of which are relevant, and is still likely to miss key papers in the field.
A good paper will cite at least 20 other sources of information. You won’t end up citing every research paper you gather.
Identify sources of data
A lot of strong papers for this world involve data. Collecting data can be hard! You might be able to run a study with an existing dataset, like the Reddit r/legaladvice data Suffolk gathered to build the Spot machine learning classifier.
Ideally your data can be shared; at least try to share a subset of it so other researchers can evaluate the quality of your results.
If you cannot find natural sources of data, think about ways to create realistic test data and annotate it, even by yourself as long as you can state with confidence that your annotation is a good proxy for an expert in the same setting. For this, being a practicing lawyer can be a superpower. In Getting in the Door, Hannes and I created realistic scenarios to test the LLM on our own, and used ChatGPT to help us add a larger variety of questions. I had plenty of experience with the realistic scenarios from 12 years of practicing as a housing attorney.
Design an experiment
You now have a research question and a source of data. How will you use the data to answer your question? What kind of results will be sufficient to answer your research question?
For example: maybe you want to show that an LLM can effectively score or answer natural language user questions. An experiment might compare the LLM’s rating to that of human experts. The more experts you have, the stronger your results, but even a quick analysis from a few peers could be interesting.
It is important to think in advance about isolating your research questions (which should be set at the beginning of your project) from the process of evaluating the experiment. You do not want the results to change your questions, making them invalid.
Consider strengthening the results by expanding the scope. For example: instead of running a study on a few examples in ChatGPT, gather hundreds of test cases and run against multiple models. In both Getting in the Door and That’s so FETCH, the paper was strengthened by evaluation across many different kinds of large language models, and doing so added only a small bit of effort.
Think through relevant metrics
You can always ask an expert, but some metrics will make more sense to you than others. Look through the list below to get a sense of what kind of statistical scores are used in papers in this field. These are not all exactly equivalent metrics, and you might include one or more in the same paper.
| Metric | What it measures | More info |
|---|---|---|
| Accuracy | % of correct predictions | Accuracy (Wiki) |
| Precision / Recall / F1 | Balance of false positives/negatives | Precision & Recall (Wiki) |
| Cohen’s Kappa | Agreement between AI and humans, adjusted for chance | Cohen’s Kappa (Wiki) |
| Intraclass Correlation Coefficient (ICC) | How strongly multiple human raters agree on continuous ratings | ICC (Wiki) |
| Krippendorff’s Alpha | Inter-rater reliability (multiple coders) | Krippendorff’s Alpha (Wiki) |
| MAE / RMSE | Error in continuous predictions | MAE & RMSE (Wiki) |
| AUC-ROC / AUC-PR | Quality of classification under different thresholds | ROC & AUC (Wiki) |
| Completion rate | % of users who finish a task (e.g. form) | Usability Metrics (NNGroup) |
| Task success rate | Whether users accomplish the intended task | ISO 9241-11 |
| Time on task | Efficiency: how long a task takes | Usability Metrics (NNGroup) |
| User error rate | % of mistakes made by users | Usability.gov |
| User satisfaction | Subjective experience (e.g. Likert scale) | System Usability Scale (SUS) |
| Disparate impact ratio | Fairness across demographic groups | Disparate Impact (Wiki) |
| Equalized odds | Classification fairness across groups | Fairness in ML (Wiki) |
| Error distribution | Which groups get more errors | Fairness in ML (Wiki) |
| BLEU / ROUGE | Text similarity to a reference (might be used when studying hallucination rate) | BLEU (Wiki), ROUGE (Wiki) |
| Human evaluation | Expert judgment on accuracy/quality | Human Eval in NLP (ACL) |
| Legal validity rate | % of outputs legally coherent | Hallucinations in LLMs (arXiv) |
| Citation accuracy | % of correct vs. fabricated citations | LLM Hallucination Issues |
| Case throughput | Clients served per unit staff time | ThomsonReuters |
| Error correction rate | % of AI outputs needing manual fixes | Error Rate Metrics (ISO 9126) |
| Adoption / retention | Continued use over time | Adoption Metrics (Wiki) |
| Downstream outcomes | Real-world legal outcomes (e.g. housing retained) | Access to Justice Metrics (Harvard) |
Run an evaluation
Getting the statistics for your paper used to be much harder. Many data scientists would turn to a Python notebook. For studies centered around the performance of a prompt to an LLM, Promptfoo is a better alternative. Dave Guarino was kind enough to write a comprehensive three part guide (part 2, part 3) to using Promptfoo, which I followed. There are some alternatives; one you might consider is GitHub models.
The basic idea in most evaluations is to find ground truth (e.g., human annotated responses) and then to compare that to the output of your tool. Promptfoo allows you to run a single prompt against hundreds of examples and a dozen different AI models in a batch and compare the answer the AI gives to your human-annotated ground truth. You can do this with a simple YAML configuration file and a spreadsheet with your test data. The most powerful thing about Promptfoo is the ease with which you can re-run an evaluation after a change to your prompt.
Once you have the results, getting nice graphs from them is also easier than ever. You can use a Python notebook or Excel or your favorite statistical modeling tool, but you can also load the data into ChatGPT and have it create the Python code to make a graph for you using its tool mode. Using the thinking mode of your LLM can improve confidence in the output. You absolutely still must check the code that it generates for you to validate that it is correct, but typically it’s just a few lines of Python that you can easily inspect. Work with an expert here if you are not confident in your ability to verify the output.
Write about your results
You’ve run your study. You’ve done the research. Now you need to put pen to paper.
- Find the conference’s templates.
- I recommend using LaTeX in Overleaf instead of using the Word template. LaTeX is more complex to start with, but Overleaf’s visual editor makes it relatively easy to use, and once you are used to it, it is more reliable.
- Don’t let this stress you out, though! If you cannot imagine learning to use a new word processor, Word is accepted at both conferences. Just be ready for more back-and-forth at publishing time if you use Word. The Word template is usually an afterthought, and it tends to have more bugs.
- Come up with an outline for the paper.
- Think through ways to illustrate your dataset and the results of your study, with charts, graphs, screenshots, and tables.
- Start writing.
- Give it a critical review and then rewrite as necessary.
- Make sure to anonymize it if necessary (ICAIL uses double-blind reviews; JURIX uses single-blind).
Overleaf is straightforward to use, but sometimes you need to write raw LaTeX code. You do not need to turn into a LaTeX expert. I use ChatGPT to help me fix the format of a broken table or to help me add a missing package to fix smart quotes, for example.
A standard paper outline
- Abstract
- Introduction and motivation
- Research questions (your hypotheses before running a study)
- Prior work (a description of past research, by you or others, and an explanation of why it is relevant to your topic).
- Background and conceptual framework (give us relevant background and context for your project)
- Data and materials
- Methodology
- Results
- Discussion of your results (including limitations and implications as well as ethical implications)
- Conclusion (connecting the conference themes with your results and pointing to future work)
- References (20-30 relevant citations is a common floor for a well-researched paper; more than 50 might indicate clutter and lack of care in choosing relevant citations, unless your paper is a review or survey.) References are not usually not part of the page limit.
You do need most of these sections for every paper, but balance the time you spend on each section so that most of your paper focuses on the substance and your novel contributions.
Avoid an unnatural academic style
There are a lot of poorly written papers in academia. Writing for these conferences can be formal, but don’t try to fake academese. Trying to imitate this voice is not going to impress the reviewer. Make your paper stand out by writing with clarity and through presentation of strong, original ideas.
Many of the reviewers for this conference have English as a second language. Using plain language skills will help you communicate clearly and effectively.
- Use active voice.
- State your points directly.
- Use topic sentences and transitions between sections.
- Define jargon and think carefully about whether you need it.
- Split each idea into its own sentence instead of writing long run-ons.
- Use clear headings, tables, and bulleted lists to format text for easy understanding.
- Avoid idioms. Puns in your title are the one exception!
- Avoid colloquial phrasing and contractions.
- Try reading your paper out loud to spot awkward phrasing.
- Consider running your paper through a vocabulary profiler. Take time to think about whether you can use simpler vocabulary words to replace uncommon words.
- Re-read your work and cut out repetitions and throat-clearing phrases.
Do not blindly accept the changes that the grammar checker in Word or Overleaf suggests. Often they make your paper too boring or can simply be wrong. But do not ignore them, either.
The more pleasant your paper is to read, the more likely it is that the reviewer will enjoy it, and the more likely it is to be useful to future researchers or practitioners. If the reviewer cannot easily understand what you write, they might not think you could understand it, either.
Be frank about the limits
The reviewers will spot the flaws in your study, and overconfidence is going to hurt your credibility. Be frank about any limits. Do not overstate your results.
Show your work
Explain the details of your analysis. If you can share code, a prompt, data, these will all make your paper more attractive and useful for future researchers. The CS community has a preference for working transparently.
Spell out the connections you want the reader to make
While a connection might be obvious to you, do not assume your reader knows what you mean or why your results are important. Be specific about what you want them to take away from your paper.
Stay within page limits
You might get an automatic “reject” if your paper is over the page limit. Take the limits seriously.
Find a critical reader (or use ChatGPT)
My favorite ChatGPT prompt while writing is “act as a critical reviewer for JURIX or ICAIL. Identify flaws and anything unclear in this paper.” I run this prompt a few times on both ChatGPT and Gemini, and I have found it to be very helpful to make sure things are clear.
A friend is also great! But depending on your circle, a friend may not spot issues with your analysis.
Go forth and come up with something great!
I hope I get to see you at the next edition of one of these conferences! I’ve found it to be deeply rewarding.
Your next opportunity is probably ICAIL in Singapore; we also hope to host another AI4A2J workshop, which might be even more friendly than the main conference track.
