<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://annefried.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://annefried.github.io/" rel="alternate" type="text/html" /><updated>2026-09-10T01:08:15-07:00</updated><id>https://annefried.github.io/feed.xml</id><title type="html">Annemarie Friedrich</title><subtitle>your description</subtitle><author><name>Annemarie Friedrich</name></author><entry><title type="html">How to defend your PhD</title><link href="https://annefried.github.io/posts/2025-12-17-phd-defense" rel="alternate" type="text/html" title="How to defend your PhD" /><published>2025-12-17T00:00:00-08:00</published><updated>2025-12-17T00:00:00-08:00</updated><id>https://annefried.github.io/posts/phd-defense</id><content type="html" xml:base="https://annefried.github.io/posts/2025-12-17-phd-defense"><![CDATA[<p><em>This relates to PhD defenses in computer science / computational linguistics in Germany! The customs of your department/country/scientific discipline may differ.</em></p>

<p>Towards the end of their PhD journey, PhD students are experts on their research topic and probably already well acquainted with how to create posters about their research papers and with how to give an effective research talk that mostly covers one topic. What they have often rarely done, however, is giving a coherent talk summarizing all of their contributions. Moreover, PhD defense talks constitute a talk genre in their own, differing from both conference presentations and typical conference keynotes or invited talks. To provide some guidances, I am summarizing some expectations and tips in this post.</p>

<h2 id="key-objectives">Key Objectives</h2>

<p>The objective of a conference paper talk is make the audience curious about your work. The objective of an invited talk or a keynote are (in my opinion) to educate and entertain. So what is the goal of a PhD defense? Do you have to specifically defend what you’ve worked on? The ideas that you proposed? Yes and no. Plus, in natural language processing (NLP) or computer science (CS) more broadly, most content of your dissertation will have been peer-reviewed before publication at conferences or in journals beforehand, so it is unlikely that this part will be questioned. The goal is somewhat broader.</p>

<ul>
  <li>
    <p><strong>Demonstrate Experience:</strong> Your goal is to prove that you have become a competent and independent researcher in your field.</p>
  </li>
  <li>
    <p><strong>Demonstrate Awareness of Research Landscape:</strong> Often, PhD defense presentations only focus on the research bits conducted by the PhD candidate. Of course, that is an important part, but it is also important to precisely locate that bit in the research landscape. Do not forget to explain what the state of the art was when you started each bit of research and what the gap was that your approach/method/idea/results solve.</p>
  </li>
  <li>
    <p><strong>Present your work:</strong> Typically, a PhD defense is between 30 and 60 minutes, so naturally, there is not sufficient time to explain everything you have worked on. Of course, you need to give an overview, but ideally, choose two publications (three if your PhD defense is a bit longer) for which you provide a deep dive.</p>
  </li>
</ul>

<h2 id="audience">Audience</h2>

<p>In contrast to the conference presentation that you have probably given, the audience of your defense typically has a somewhat broader scientific background. This means that your presentation should explain everything in a way that at least every committee member can make sense of it. The examiners of your written thesis will have read your dissertation in detail. However, there are likely additional examiners for the oral exam which may not have read the dissertation in detail. Plus, you also want to make the rest of the audience happy. So explain everything in accessible and general terms, but be specific and go a bit more into detail when you deep-dive into those 2-3 papers.</p>

<h2 id="structure">Structure</h2>

<ul>
  <li>Start with explanations of the topic that you trying to solve, and the respective state-of-the-art.</li>
  <li>What is the motivation for working on this topic? Which problems will it solve? Think broadly. What are relationships to established subfields or research areas in NLP/CL?</li>
  <li>Create a slide that visualizes which topics/problems you have worked on and how they relate. In the image below, you can see examples from my own defense and from Stefan Grünewald’s (copyright for his slide: him and Bosch).</li>
</ul>

<p><img src="https://github.com/annefried/annefried.github.io/blob/master/images/defense/phd-defense-structure.png?raw=true" alt="example slides with structure of PhD, how topics related in boxes and in a triangle, and citations of papers" /></p>

<ul>
  <li>Mention the publications explicitly, drop the citations (in a unified short format) whenever they fit. I typically use blue color for own citations and black for other relevant citations.</li>
  <li>Most important, you need to convey what your <strong>contributions</strong> are. You know what counts as contributions because you already listed them explicitly in the introduction sections of your papers and in your PhD dissertation. This is analogous. Make it more entertaining than showing a list. The contributions do not necessarily have to occur as one list on a slide. You can explain them along the presentation whenever they fit and summarize in the end.</li>
  <li>Use (animated) diagrams to explain your methods and results (as usual).</li>
  <li>Describe implications of your work and which opportunities for future work you see, in particular those opened up by your research.</li>
</ul>

<h2 id="preparation-time-management-feedback-and-dry-runs">Preparation, Time Management, Feedback, and Dry Runs</h2>
<p>Making PhD defense slides takes time. Plan for around 2 weeks structuring and editing time, and make sure you have appointments with people helping you. Have a dry run with your team or with experienced researchers, e.g., PostDocs at your department.
Personally, I am highly grateful to my PhD advisor Prof. Dr. Manfred Pinkal and to Dr. Stefan Thater, whose last-minute feedback was extremely helpful and resulted in an evening of last-minute updates on my side but also a much improved presentation!</p>

<p>Attend as may PhD defenses at your department as you can! See what you like/dislike, what went well, what did not go so well. Plus, there is often free food. :)</p>

<h2 id="discussion">Discussion</h2>
<p>A large part of the PhD defense is typically the discussion. There is absolutely no need to be afraid of the discussion. You are the absolute expert about your topic in the room, and you have just spent at least three years engaging actively in your research community. You know what is going on! Demonstrate your reflection about the research area by discussing with other experts what you think matters. Be aware that some questions may be a bit more high-level and test your ability to generalize and draw connections also on the meta-level. Answer both in a general and in a specific way! You will be surprised how easy and interesting the discussion will feel. Often, only the members of the committee may ask questions; sometimes also everyone in the room who is a PD (Privatdozent) or professor.</p>

<h2 id="celebration-">Celebration !!</h2>
<p>I was lucky enough to do my PhD at the CoLi department of Saarland University (now called Language Science and Technology). This department really celebrated PhD defenses. Everyone went. There was a formal handshake and congratulations, and drinks and food afterwards. The fellow PhD students always created a hat with silly items on it. The recent graduate has to guess what the items stand for, and only if they guessed everyone correctly, they were allowed to put on the hat. It was an awesome time, and worthy celebration of x years of hard work!</p>

<p><img src="https://github.com/annefried/annefried.github.io/blob/master/images/defense/phd-hat.png?raw=true" alt="three pictures of the social part of my PhD defense with the hat" /></p>

<p>In the pictures, you can see</p>
<ol>
  <li>me celebrating with my PhD advisor Manfred Pinkal, a linguist and computer scientist and pioneer in computational linguistics</li>
  <li>me guessing what the items on the hat mean</li>
  <li>me having guessed everything correctly and wearing the hat</li>
</ol>

<p><em>Enjoy your PhD defense!</em></p>]]></content><author><name>Annemarie Friedrich</name></author><category term="scientific presentations" /><category term="nlp" /><category term="phd" /><summary type="html"><![CDATA[This relates to PhD defenses in computer science / computational linguistics in Germany! The customs of your department/country/scientific discipline may differ.]]></summary></entry><entry><title type="html">Conference Report ACL 2025 - How can we move forward as a field?</title><link href="https://annefried.github.io/posts/2025-08-02" rel="alternate" type="text/html" title="Conference Report ACL 2025 - How can we move forward as a field?" /><published>2025-08-02T00:00:00-07:00</published><updated>2025-08-02T00:00:00-07:00</updated><id>https://annefried.github.io/posts/conference-report-acl2025</id><content type="html" xml:base="https://annefried.github.io/posts/2025-08-02"><![CDATA[<p><i>DISCLAIMER: If you disagree with me, or if I represented your opinions wrongly in this article, please contact me directly via e-mail! I am happy to adapt this article if needed.</i></p>

<p>Last week, I attended the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025). My PhD advisor Manfred Pinkal told me recently that the first ACL he attended featured 150 attendees. This year, there were more than 5000 onsite and roughly 1000 additional remote participants. It was also the first time I attended a huge conference after the Covid break, so it was extra nice to catch up with a lot of people in person.</p>

<p>Every invited talk and panel (except maybe that by Mark Johnson in the XLLM workshop) that I attended spent the first ten minutes discussing the great capabilities of large language models (LLMs) and why they do not yet solve language understanding, or some related tasks. Moreover, ACL ARR has a massive problem with way too many submissions, too few senior reviewers, and too many automatically generated reviews. In a plenary <a href="https://2025.aclweb.org/program/panel/">panel discussion</a>, Ed Hovy, Mirella Lapata, Yue Zhang, and Dan Roth discussed how we could move forward as a field.</p>

<ul>
  <li>
    <p><strong>Ed:</strong> With LLMs, we are given a self-driving car that just crashes regularly. We should spend more effort on understanding how it works internally. I think this is really worthwhile, though he really meant that we should develop methods that let us understand the internal workings of the model such that we can fix ‘misbehavior’. All the papers that just probe the models on some task and then say, hey, LLMs can do this or that (to some degree) are like popcorn (his words, not mine), you read them, you feel satisfied for ten minutes, and then you feel empty again. I chatted with Ed later and (if I did not misunderstand him) he actually does agree that system-building using LLMs as a component is still a worthwhile endeavor, if that’s your focus. I will come back to this point later. As a community, we have tried understanding how neural models work for quite a while, e.g., along the lines of the <a href="https://blackboxnlp.github.io/2025/">BlackBoxNLP workshop series</a>, which is still active. There were some interesting insights, e.g., that <a href="https://aclanthology.org/P19-1452/">BERT behaves a bit like the old natural language processing (NLP) pipelines</a> with higher-level semantics being put together at higher levels. However, I think we have failed to even understand simple models like word2vec fully in the sense of being able to predict when they succeed and when they fail. In my current opinion (which I am happy to change if someone brings evidence or good arguments), LLMs are still distributional models, with just highly sophisticated vector spaces that enable us to store a lot of information how to traverse the vector space and output natural language answers at the end. It does make sense to look at where they succeed and where they fail. In particular when building systems for real-world use, it is paramount to test them extensively and make sure they do not harm anyone. Unfortunately, the popcorn papers are overwritten as soon as the next model generation comes out. Then we have to test everything again, and unless we can be sure that the new models were not simply trained on the probing test data, we even cannot re-use the same benchmark data. We’ve already seen this with all the BERTology papers.</p>
  </li>
  <li>
    <p><strong>Mirella:</strong> We need to get back to an actual science, where we have a controlled setting of what went into the training data and on what data we test. I completely agree. This was mainly a call to the model vendors. The issue is that indexing and providing search capabilities for the pretraining data is of course costly. Plus, we would have to keep developing methods to find the instances of data that are of interest to the model behavior we aim to analyze, which may also raise new questions how to do this science. (Though I personally think doing research on such retrieval methods could be fun.) During the XLLM workshop, I learned about some interesting endeavors in this direction like <a href="https://openeurollm.eu/">OpenEuroLLM</a>. <a href="https://github.com/allenai/OLMo">OLMo</a> is also an exciting initiative, which also features an <a href="https://huggingface.co/allenai/OLMo-2-0325-32B-Instruct">instruction-tuned variant</a> (unfortunately at the moment predominantly for English). However, for most of use, regular researchers at a regular unversity, training such models is not something we can actually afford. So what can we do to make the field more scientific again? I will present my thoughts on that below.</p>
  </li>
  <li>
    <p><strong>Dan:</strong> The topic of the panel discussion was actually supposed to be “generalisation,” even if a large part of the discussion was more general than that (pun intended). Dan reported that he had asked some LLM about himself recently and got an impressive list of prizes he had won, just that he had not actually won them. But he pointed out that the model actually <i>did</i> generalize by hallucinating these facts, as they typically fit for a university professor at his stage. If I understood correctly, Dan, who was presented as the “pragmatist,” largely took the perspective that we should build and evaluate systems. But also that we should aim to build systems (or understand LLMs in that sense) how they deal with reasoning chains or, as they put it, causal reasoning. In our own work on <a href="https://aclanthology.org/2024.emnlp-main.153/">quantifying uncertainty in natural language text in Bayesian reasoning scenarios</a> (EMNLP 2024), we actually found that they perform okay on causal reasoning (inferring the likelihood of effects based on some cause), but that their performance drops markedly when the underlying problem requires evidential (updating one’s beliefs about causes when new evidence comes in) or explaining-away style reasoning (in which the knowledge about one cause makes the beliefs that some other cause is the case less likely in light of evidence). This, I think, is actually somewhat intuitive given their autoregressive nature.</p>
  </li>
</ul>

<p><img src="https://github.com/annefried/annefried.github.io/blob/master/images/quite-results.png?raw=true" alt="Results taken from our EMNLP 2024 paper: the models' performance drops for evidential and explaining-away reasoning" /></p>

<p>What is worth noting is that despite the teaser image, our paper is not actually popcorn (I hope), because our point was actually to create a neurosymbolic model that parses problems into a machine-readable logic programming language that can then solve the problems regardless of the underlying reasoning types (see blue bars in the plot). Results like this make me somehow believe that without a major change in model architecture, the models have some built-in bias that just works unlike the human brain, and that to achieve models that explain reasoning in a way that we humans can deal with, we need at least one additional architecture shift in AI. But that is a belief, and as a researcher, I occasionally update my beliefs. Evidential, causal, or explaining-away reasoning included.</p>

<p>So still, ACL has a massive problem of somehow having drifted away from the scientific principles guiding NLP research in the past decades. Getting papers accepted seems to have become some kind of gamble which highly depends on being assigned a responsible meta-reviewer. In his keynote at the Scientific Document Processing Workshop, Ed Hovy again provocatively stated that <a href="https://aclanthology.org/2024.acl-long.547/">automatic meta-reviewing had been solved</a> (don’t get me wrong, I loved his provocative analogies, our community needs leaders like him that have thought about NLP and meaning for decades). Again, in our coffee break chat, he agreed that neiter reviewing nor meta-reviewing should be summarization. Practically, the meta-reviews that my students receive regularly read like superficial summaries of the reviews. There is no meaningful evaluation of the contributions as in prior times. But it shouldn’t be that way. Ed actually thinks we should admit all the papers and vote on-site which papers should get into the proceedings. A little like it is actually the case in linguistics - in this field, contributions are often accepted for presentation at a conference based on an abstract, to check the topical fit. After discussion and feedback, editors that actually edit compile a collection of articles into a book. Maybe that wouldn’t be a bad idea, I am just not sure whether it scales with the number of ACL participants. But a lot of people I talked to at ACL this year (including also Alexander Koller) and also ACL president Chengqing Zong in his presidential address advocated going back to smaller and specialized conferences.</p>

<p>What bothers me, personally, is <strong>understanding how we should do science these days</strong>. And how to communicate this to the newcomers in our field. On my train ride back home, I was thinking about how I communicate this to my PhD students and decided to put together some concrete suggestions that can help people who do not own huge compute centers. But first, let’s look back. What were experimental setups, what were “interesting” or “valid” contributions?</p>

<h2 id="computational-linguistics-before-the-year-2000">Computational Linguistics before the year 2000</h2>

<p>For decades, computational linguistics dealt with building programs (typically not yet what we would call a system today) that would process natural language, often motivated by theoretically established linguistic rules. The contribution was as much on the linguistic side as on the computational side, as by formalizing and testing linguistic ideas in a computational way was still rather novel. For example, look at the <a href="https://aclanthology.org/P87-1003.pdf">PUNDIT system</a> for temporal relation inference (which is more on the theoretical side) or a <a href="https://aclanthology.org/M91-1023.pdf">Text Interpretation System</a> for MUC-3.</p>

<p><img src="https://github.com/annefried/annefried.github.io/blob/master/images/examples-cl-1990.png?raw=true" alt="screenshots taken from the two papers mentioned in the text that illustrate how the grammar rules look like" /></p>

<p>These papers do not even have an experimental or evaluation section. <strong>Reviewers would have to look at the ideas and judge whether they would be interesting to discuss, whether they have the potential to spark new ideas in others.</strong> Whether the approach does something different than existing work. <strong>Proposal #1: Let’s add this back to our criteria of reviewing and meta-reviewing, and not just on guidelines pages, but actually do that.</strong> It will result in much more interesting contributions to ACL than benchmark chasing. I think this fits in nicely with the necessity to create less compute-intensive models. And let’s not just do that in specific environmental computing tracks. Let’s invite diversity back into our approaches.</p>

<p>To illustrate what is currently wrong: The only paper from my group that did not get into ACL or Findings actually presented a super interesting way to create training data for a complicated semantic parsing task and showed clear improvements on medium-sized models. It was rejected mainly for the reason that GPT4o and Deepseek achieved around 61% accuracy on the dataset, and our approach only 57%. Are these really numbers that already tell us that we are on the wrong track? What if our model had achieved 63% accuracy? I think if any SOTA LLMs achieve anything less than 98% accuracy on a (non-subjective) task, it is absolutely worth discussing alternative approaches! (This paragraph is not intended to be a personal complaint, it just intends to illustrate what I also hear from lots of colleagues, despite the ACL ARR guidelines already discouraging such simplified ways of evaluating the perceived utility of an idea.)</p>

<h2 id="natural-language-processing-2000-2022">Natural Language Processing 2000-2022</h2>

<p>In the 2000s, the mainstream paradigm in NLP was to collect data, annotate it manually, compare whether the task is meaningful by at least checking if several humans would arrive at the same conclusion if they perform it, then train a machine learning model guided by their intuition and a development set not included in the training data. Finally, we would test whether the model generalizes either in- or cross-domain by applying it to some also manually-labeled test data. One could argue that often, the data sets were way too small for drawing statistically significant conclusions (although they were often much bigger than nowadays). Often, it can also be shown that the <a href="https://aclanthology.org/N18-2017/">models actually use superficial cues rather than actual understanding</a>. So I really do not want to say “everything used to be better”.
However, it was paramount to separate the data on which the model was trained from the data on which hypotheses were tested. Also, the data should either have been elicited throug human annotation (more often the case, example: <a href="https://aclanthology.org/P16-1166/">my own PhD work on linguistic aspect</a>) or they should have come from the real world (less often the case, examples from the patent NLP domain: <a href="https://aclanthology.org/2022.emnlp-main.791/">a dataset on patent classification created by a Bosch employee over the course of 25 years</a> or <a href="https://aclanthology.org/2025.findings-acl.496/">PAP2PAT, which aligns patents and papers</a>).</p>

<p><img src="https://github.com/annefried/annefried.github.io/blob/master/images/experimental-paradigm-nlp-2000.png?raw=true" alt="image summarizing the experimental paradigm 2000-2022 that is also explained in the main text" /></p>

<p>As a side note, it was also considered bad practice to <strong>reverse-engineer</strong> other systems. For example, if you take an existing dependency parser and run it on a dataset, which you then use to train another parser, you have not actually learned anything new, you have just learned the set of rules already explicitly or implicitly incorporated in the initial parser. The only admissible setting that I can think of is when the aim is to construct (typically additional) <strong>silver data</strong> that is combined with existing <i>gold data</i> and the aim is not to make statements about the underlying phenomenon that you are trying to capture but instead to improve downstream performance on an unseen test set. (I am mentioning this because I noted recently that a lot of novice researchers have not heard of the concept of silver standard data, although much of the currently LLM-created benchmark data falls in this category in my opinion. It was also suggested by Jessy Li in her keynote at the Linguistic Annotation Workshop that we should increasingly work with <i>platinum data</i>, i.e., manually created data collected by actual expert annotators.)
<strong>Synthetic data generated by an LLM is necessarily just silver data</strong>, hence, we should always take conclusions made based on it (especially if it is used as test data) with a grain of salt.</p>

<p>But what can we learn from this age? <strong>Only if the training data does not contain the test set, we can actually make statements about whether a model has learned anything beyond memorizing some instances.</strong> But how can we tell what data LLMs have seen during pre-training? Sadly, at the moment, for the most widely-used LLMs, we cannot. But we can do something. <strong>Proposal #2: We should use test data that was created newly, either by hand-labeling data (by making use of trustworthy annotations, perhaps even having them onsite – I would not trust crowd-sourcing any more) or by collecting real-world data that was definitely created after the cut-off date of the pretraining data of the LLMs that are part of the system we are testing.</strong> An analysis of such non-contaminated test data should be a mandatory part of any paper using the standard machine-learning benchmarking paradigm described above. And reviewers should give credit for it, because it is a lot of extra work!</p>

<p>Here is an example taken from the PAP2PAT paper. I really liked that my student Valentin collected this additional data and performed experiments on it, actually observing some interesting deviations on the linguistic style, and I think it would be very worth diving into such observations with an even deeper analysis. (This is not criticism of his careful work, which took more than one year, and which I, in my completely unbiased opinion, think it would absolutely have been worthy of being accepted at the main conference. :))</p>

<p><img src="https://github.com/annefried/annefried.github.io/blob/master/images/pap2pat-example.png?raw=true" alt="a screenshot of the Pat2Pat paper sections on test data contamination" /></p>

<h2 id="reflection-on-research-practices-today">Reflection on “research” practices today</h2>

<p>So why are people like Mirella Lapata saying that our research field needs to get more scientific again? The problem is that LLMs have become so human-like in their way of expression that humans easily trust them potentially too much. Steven Bird says (in <a href="https://aclanthology.org/2024.acl-long.797/">this recommended article</a>) that this is due to the Eliza Effect, the linguistic correlate of pareidolia (the human tendency to see faces in the clouds, i.e., to assume meaning where there is none). If you think about it, it does make sense: we rely on recognising other humans, other intelligent beings, so it is better to overgeneralize in recognising them by their behavior rather than to miss out on noticing one of these dangerous predators in our surroundings. Another issue that is discussed a lot these days is whether we potentially anthromorphize LLMs too much. I have several times talked to people outside our field who were suprised that “agents” are just systems using the same underlying LLM technology. They were somehow of the impression that “agentic AI” was really something fundamentally new.</p>

<p>Granted, on some tasks (for some languages), LLMs even work more consistently than trained human expert annotators. However, I do not think this is generally true for all relevant setups.
Moreover, saying that they have <a href="https://aclanthology.org/2023.acl-long.697/">superhuman performance is really misleading</a>, and I think that is still true. At the very least, we should define human performance with regard to that person who is best at performing a specific task, and not with regard to an average crowdworker.</p>

<p>But psychology set aside, what is the issue? A lot of papers (including those I co-authored, I am sure you can find examples of what should be done differently in those as well!) <strong>generate training or test data</strong>. In some domains, this is extremely helpful as it may be impossible to work with real data, e.g., due to copyright or privacy issues. So in some sense, the possibility to actually generate new texts or data is highly beneficial to our research. There is only one issue: the data already follows some internal distributions of the model(s). Depending on the use case, this may not be a big problem. If you intend to build a useful system, it may be totally okay, <strong>as long as you test on real-world data</strong>. Often, a test dataset is simply built in the same way, which means that we may overestimate performance.</p>

<p><img src="https://github.com/annefried/annefried.github.io/blob/master/images/cyclic-research-issues.png?raw=true" alt="image of a cycle going through the steps described in the main text, with issues" /></p>

<p>I have seen some papers (during the review process) that then use the LLM to label the data, and perhaps also additionally filter the data “to increase data quality.” Maybe that increases the quality of the instances that were labeled, but it certainly also discards the relevant cases. And getting the hard cases along the decision boundaries right should remain the core interest of NLP, not dealing with test sets whose instances are at the far ends of the spectrum.</p>

<p>Then, the “modeling” takes place. I think you see where I am going - if the same LLM is used to model the data that we just automatically generated and/or labeled and filtered, it has an unfair advantage because it uses the same underlying distributions. If you use a different LLM, you risk reverse-engineering the data-generating model. If they have not already been aligned during pre-training and instruction tuning. In my personal experience (does anyone know a good citation for this?), very different models sometimes tend to make exactly the same mistakes on difficult cases, which lets me wonder to what extent the training data for one model was actually taken from the other model. Sigh. This is the opposite of scientific.</p>

<p>Finally, evaluation. Real-world data with reliable annotations is extremely hard to get by. During my PhD, I worked with a group of several trained student annotators for more than three years. Who does something like that any more nowadays? The next ACL ARR cycle is due in two months. Luckily, we have LLMs that behave almost like humans, so that is why the LLM-as-a-judge (a special case of LLM-as-an-annotator according to Rotem Dror, see her keynote in the Linguistic Annotation Workshop), became popular. LLMs are still incredibly unreliable at outputting the solutions in the format we want to have it, i.e., we often need to map it to other formats or find out if it contains the correct answer (and no incorrect one). Plus, it is still tricky to evaluate the quality of generated text. It is easy to see that when using an LLM that was used to generate or label the data, it will prefer its own labels. There exists already a growing body of work on the biases of models (and humans) on judgment tasks, revealing systematic differences.</p>

<p><img src="https://github.com/annefried/annefried.github.io/blob/master/images/cyclic-research-important-points.png?raw=true" alt="image of a cycle going through the steps described in the main text, with potential solutions" /></p>

<p><strong>Okay, but what can we do (to do scientific research)?</strong> For a moment, let’s stay in the machine-learning style NLP framework and see how we can improve each step to break the cyclic research.
Here is my proposal.</p>

<ul>
  <li>
    <p>Let’s <strong>base our research on open models for which we can actually verify whether something was included in the training data or not</strong>. Even if they are smaller, even if Deepseek, ChatGPT, or LLama-3.3-405B (replace this with the currently biggest commercial and open-weights models of your choice) perform better on some benchmark. This should really, really be not a factor in reviewing.</p>
  </li>
  <li>
    <p>If we are using LLMs to <strong>label data, verify the quality of this step by comparing to real-world data (that had also not been included in its training) or human expert annotations</strong>. There is a lot of recent work by Rotem Dror on statistical significance tests for agreement. I also recommend studying <a href="https://aclanthology.org/J08-4004">this article</a> by Ron Artstein and Massimo Poesio. It’s from 2008, but the metrics are still widely used, account for chance agreement, and if you look at them class-wise, they can tell you which are the hard and easy cases. We can simply not say generally how many instances we need to annotate manually to build a good labeling prompt. <a href="https://aclanthology.org/2023.eacl-main.38/">Class imbalance</a> has a major impact on performance of both humans and systems, and there are simply easier and <a href="https://aclanthology.org/P14-2064/">difficult cases</a>. During the crowdsourcing age, a lot of methods have been developed to figure our statistically which annotations to trust, e.g., Dirk Hovy’s <a href="https://aclanthology.org/N13-1132/">MACE</a>. I think these methods could be quite useful to figure out whether to trust LLM annotations as well.
To study agreement, I invite everyone to look into the old literature. :)</p>
  </li>
  <li>
    <p>From a system building perspective, I think we would kind of be done at that step. We have collected a small, but meaningful set of annotations, written some rules (or programmed the LLM with some prompts if you prefer this terminology) and checked with the data how well our rules capture the quality annotations. That was exactly how NLP was done for a long time, and it actually does not seem wrong to me these days either. I would like to read about this step of the research much more. <strong>The problem is (as it was before) that if the person who wrote the rules is also the annotator, the performance measure may be overly optimistic again.</strong> While I studied during my Master’s degree in Language Science and Technology in 2010, we were made well aware of this issue.</p>
  </li>
  <li>
    <p>For historical reasons, after the data collection, the modeling happens. Now some additional prompts are designed, perhaps an agent system is put to use. If your system really does more than the original labeling prompt (maybe you finetuned a model and want to show that it made progress by solving a particular task), it should next be evaluated. As before, I think <strong>when using an LLM-judge, it is paramount to estimate the quality of the judge on the particular task</strong>. We should always, always, always do that unless the particular judge has already been evaluated for the exactly same test data by a predecessor paper. And then we should account for the variance introduced by the potential uncertainty of the judge when we compare systems. Not sure how to do this statistically, I am sure this would be a great research topic for you, Rotem Dror, let me know if you are interested in collaborating if you read this. :)</p>
  </li>
</ul>

<h2 id="looking-sideways-research-methods-in-hciir">Looking sideways: research methods in HCI/IR</h2>

<p>In my opinion, a lot of work feels less scientific because the methodology conflates experimental paradigms from machine learning and from system building. I think we need to more clearly position research papers which experimental paradigm they follow. As I have explained above, the train-dev-test data paradigm was not always predominant in computational linguistics. So I think while it is important to go back to it for research work that aims at improving the machine learning step itself, we should also acknowledge that part of our research today kind of went back to system building. For machine learning experiments, we should rely on transparent models and obtain the kind scientific trustworthiness back that Mirella is advocating. As for approaches that build systems, we should just not try to sell it under a wrong paradigm and use evaluation methodology that employs LLMs as questionable judges.</p>

<p><strong>Let us consider again for whom we are building language technology.</strong></p>

<ol>
  <li>
    <p>Other humans. Humans work differently from LLMs. Even if LLMs-as-a-judge achieve good correlations with human responses, I advocate that systems should still be evaluated by the distribution of human users that are the intended users. This is pretty common in other scientific research fields such as human computer interaction (HCI) or information retrieval (IR). It does not always have to be a questionnaire, we can also track gaze, click traces, etc. (with the users’ consent of course).</p>
  </li>
  <li>
    <p>Other systems. The output of our NLP approach may be entered into databases or fed into business processes, etc. For each step, there should also be human experts that should inspect several cases. They should be sampled in a way that they cover the entire range of rare, frequent, difficult, and easy cases. Potentially, this step can also simply be evaluated based on real-world data.</p>
  </li>
</ol>

<p>What does this mean? <strong>Authors should clarify under which experimental paradigm they work, and reviewers should not discount research that makes its claim based on user studies rather than F1 scores.</strong> Maybe that would help to avoid papers that use questionable research designs.</p>

<h2 id="summary-how-can-we-move-forwards">Summary: how can we move forwards?</h2>

<ul>
  <li>
    <p>Any scientific hypothesis can only be proven if we KNOW that a model has not seen the test data during pre-training or instruction tuning. If a paper uses open models (which I hope will become more common during the next year) and/or describes how they checked for test data contaminations, this should be credited strongly positively in reviews and meta-reviews.</p>
  </li>
  <li>
    <p>Again, please don’t get me wrong. I think it is okay to use LLMs for system building. We should just distinguish carefully between <strong>research</strong> on systems and <strong>engineering</strong> systems. While engineering can be research, too, in the context of LLMs it helps to think about the fact that submitting a patent application that just says “my invention is to perform task X with an LLM” will highly likely be rejected. By contrast, if a research paper on a system performs an insightful analyses on the capabilities, boundaries, component statistics, and potential impact of the system (ideally by testing it using the actual target users), that could count as research. As a research community, we should also perform research on the current capabilities of models that are considered to be on the forefront of our field. Perhaps even popcorn papers are important to raise awareness of all the sociological and ethical problems that come with the spread of LLMs. But precisely for that reason, humans should be kept in the loop when evaluating such models or systems.</p>
  </li>
  <li>
    <p>LLMs are already actively in use to support research in other fields. And not only to polish texts or search the literature, but also to obtain data from this (according to Luke Zettlemoyer) huge mountain of compressed data. Which may not be wrong per se, but unless we carefully test the data that we obtain from it, we risk drawing misleading conclusions from the synthetic data. I think it is paramount that we as NLP researchers set a good example how to deal with this data.</p>
  </li>
</ul>

<p><strong>Is there AI-generated content in this article?</strong> I did not use any AI to write this article. Like Iryna Gurevych, who pointed this out in response to Ed’s talk at SDProc, I still find it easier and faster to directly find the words that I can use to express something rather than finding the prompt that generates the text that I actually want. Maybe that is because I was actually trained extensively in writing, and my way of doing things is old-fashioned. Like someone who still remembers landline numbers rather than searching for a phone number in their Google contacts. Sometimes the old way is just faster. (I have to admit that I have to look up even my husband’s cell number these days, though.)
I do not know whether this means that we should be worried as a generation of researchers and developers that learn neither to write nor to program themselves is in a critical education stage these days. Maybe I am really just old-fashioned. On the other hand, how could I learn to write prompts for a translation system and judge if the output is what I want if I do not actually speak the target language?
Maybe I could smoothe the language a bit (I am not a native English speaker). However, the article is authentic, the words are mine, not those of some LLM.</p>

<p><strong>Thanks</strong> to Casey Kennington (Boise State University) for finding my typos and suggesting to add or elaborate on several interesting aspects.</p>]]></content><author><name>Annemarie Friedrich</name></author><summary type="html"><![CDATA[DISCLAIMER: If you disagree with me, or if I represented your opinions wrongly in this article, please contact me directly via e-mail! I am happy to adapt this article if needed.]]></summary></entry><entry><title type="html">Natural Language Processing Datasets in the Materials Science Domain</title><link href="https://annefried.github.io/posts/2025-11-08-materials-science-datasets" rel="alternate" type="text/html" title="Natural Language Processing Datasets in the Materials Science Domain" /><published>2024-11-08T00:00:00-08:00</published><updated>2024-11-08T00:00:00-08:00</updated><id>https://annefried.github.io/posts/materials-science-datasets%20copy</id><content type="html" xml:base="https://annefried.github.io/posts/2025-11-08-materials-science-datasets"><![CDATA[<p><strong>WORK IN PROGRESS</strong></p>

<p><a href="https://en.wikipedia.org/wiki/Materials_science">Materials science</a> is an interdisciplinary field focused on the research and discovery of new materials. During my time at Bosch, I worked closely with several domain experts from this field and learned to appreciate how important it is to identify materials with particular characteristics in order to solve challenges such as addressing climate change (e.g., by building environment-friendly energy systems or cars). But this blogpost is not about the benefits of materials science research, it is about how natural language processing (NLP) can support this research.
While much prior work in NLP for scientific text has concentrated on the biomedical domain, there is a growing interest in also building datasets and solutions for the materials science domain.
In this blogpost, I am collecting corpora and benchmarks addressing this domain, which are the basis for building solutions that can meaningfully help materials science researchers.</p>

<p><em>If you have / know a dataset that falls under this category that is not listed here yet, please just e-mail me!</em></p>

<table>
  <tbody>
    <tr>
      <td>Dataset</td>
      <td>Reference</td>
      <td>Data / Annotations</td>
      <td>Materials science subdomains</td>
    </tr>
    <tr>
      <td>SOFC-Exp</td>
      <td>…</td>
      <td> </td>
      <td> </td>
    </tr>
  </tbody>
</table>

<h2 id="sofc-exp">SOFC-Exp</h2>
<p>https://aclanthology.org/2020.acl-main.116.pdf</p>

<h2 id="materials-science-procedural-text-corpus-mspt">Materials Science Procedural Text Corpus (MSPT)</h2>
<p>https://aclanthology.org/W19-4007/</p>

<h2 id="recipes">Recipes</h2>
<p>https://www.nature.com/articles/s41597-019-0224-1</p>

<h2 id="weston--ner">Weston / NER</h2>
<p>https://pubs.acs.org/doi/10.1021/acs.jcim.9b00470</p>

<h2 id="mulms">MuLMS</h2>
<p>Text: https://github.com/boschresearch/mulms-wiesp2023
https://github.com/boschresearch/sciol-wacv-2024
https://github.com/boschresearch/mulms-az-codi2023</p>

<p>https://openaccess.thecvf.com/content/WACV2024/papers/Tarsi_SciOL_and_MuLMS-Img_Introducing_a_Large-Scale_Multimodal_Scientific_Dataset_and_WACV_2024_paper.pdf</p>

<h2 id="polynere">PolyNERE</h2>
<p>https://aclanthology.org/2024.lrec-main.1126.pdf</p>

<p>maybe: https://aclanthology.org/2024.findings-acl.779.pdf</p>

<h2 id="ms-mentions">MS-Mentions</h2>
<p>https://aclanthology.org/2021.emnlp-main.101/</p>

<h2 id="sc-comics">SC-CoMIcs</h2>
<p>https://aclanthology.org/2020.lrec-1.834.pdf</p>

<h2 id="polyie">PolyIE</h2>

<h2 id="pcmsp">PcMSP</h2>
<p>https://aclanthology.org/2022.findings-emnlp.446/</p>

<h2 id="other-related-works">Other related works</h2>
<p>https://aclanthology.org/2021.emnlp-main.438.pdf
https://aclanthology.org/2023.acl-long.753.pdf</p>

<p>They used data from several datasets!
https://aclanthology.org/2023.acl-long.201.pdf</p>]]></content><author><name>Annemarie Friedrich</name></author><summary type="html"><![CDATA[WORK IN PROGRESS]]></summary></entry><entry><title type="html">Bachelor/Master Thesis Opportunity: Detecting Reported Speech with Large Language Models</title><link href="https://annefried.github.io/posts/2024-08-23-thesis_reported_speech" rel="alternate" type="text/html" title="Bachelor/Master Thesis Opportunity: Detecting Reported Speech with Large Language Models" /><published>2024-08-23T00:00:00-07:00</published><updated>2024-08-23T00:00:00-07:00</updated><id>https://annefried.github.io/posts/thesis_reported_speech</id><content type="html" xml:base="https://annefried.github.io/posts/2024-08-23-thesis_reported_speech"><![CDATA[<p><em>Position filled!</em></p>

<p>Are you fascinated by the crossroads of machine learning and human culture? Work with us on a Bachelor/Master Thesis that delves into finding out who said what to whom!</p>

<h2 id="the-challenge">The Challenge</h2>
<p>You will develop an algorithm based on state-of-the-art deep learning technology and make use of large language models that detect parts of texts that report something originally said by someone.
Ideally, your solution will also attribute the said information to its speaker. The special challenge is to make your solution not only work for English, but also for German text (training and evaluation data is available).</p>

<h2 id="interdisciplinary-collaboration">Interdisciplinary Collaboration</h2>
<p>Bridge the gap between computer science and humanities by creating models that resonate with domain experts, fostering collaborative research. One use case addressed in this thesis will be addressed together with the Chair of German Linguistics (Prof. Dr. Sonja Zeman) at the Department of Philology and History.</p>

<h2 id="why-choose-this-thesis">Why Choose This Thesis?</h2>
<p><strong>Innovation:</strong> Apply cutting-edge machine learning to real-world challenges in Digital Humanities. You will get practical experience with pre-trained large language models / foundation models and deep learning frameworks such as HuggingFace/PyTorch.</p>

<p><strong>Impact:</strong> Contribute to the advancement of knowledge preservation, cultural analysis, and interdisciplinary collaboration.</p>

<p><strong>Guidance:</strong> During your thesis, we are committed to providing excellent and regular advice on experimental design, implementation, and scientific writing. Work with and get mentorship from experts in natural language processing and linguistics.</p>

<h2 id="prerequisites">Prerequisites</h2>
<p>Open to passionate Bachelor’s/Master’s students with an interest in natural language processing, machine learning, strong Python skills, and a drive to make a meaningful impact.</p>

<h2 id="contact">Contact</h2>
<p>Prof. Dr. Annemarie Friedrich / Dr. Jakob Prange</p>

<p><a href="https://www.uni-augsburg.de/de/fakultaet/fai/informatik/prof/coling/">https://www.uni-augsburg.de/de/fakultaet/fai/informatik/prof/coling/</a></p>]]></content><author><name>Annemarie Friedrich</name></author><category term="thesis_opportunity" /><summary type="html"><![CDATA[Position filled!]]></summary></entry><entry><title type="html">Scientific Writing</title><link href="https://annefried.github.io/posts/2020-09-01-scientific_writing" rel="alternate" type="text/html" title="Scientific Writing" /><published>2024-05-27T00:00:00-07:00</published><updated>2024-05-27T00:00:00-07:00</updated><id>https://annefried.github.io/posts/scientific_writing</id><content type="html" xml:base="https://annefried.github.io/posts/2020-09-01-scientific_writing"><![CDATA[<p>This blog post summarizes some important hints and tricks when writing a scientific paper or a thesis in the area of natural language processing (NLP) or computational linguistics (CL).
Of course, they are only recommendations and the ideal solution depends on your concrete topic.</p>

<p><strong>When you are advised by me or a team member on your thesis, please work through this page carefully. Feel free to ask any remaining questions about your specific thesis structure in one of our meetings. Please submit a draft for a chapter early (at the latest midway into your thesis time) for detailed feedback.</strong></p>

<p>Some initial hints:</p>

<ul>
  <li>First, very importantly, <strong>scientific writing is not like crime story writing</strong>: you do not need to wait for your results until the end. Throughout your article or thesis, mention important points and findings early on.
This will help the reader to understand and contextualize your proposed method/dataset/findings, and if your results are cool, motivate them to keep reading.</li>
  <li><strong>Write top-down</strong> rather than bottom-up: each paragraph should start with a sentence that makes the reader expect what is going to be explained. Each section should start with 2-3 sentences stating what is going to be explained in the section. Each chapter … (I think you got the idea). The Introduction should tell the reader what is happening in the article/thesis!</li>
  <li><strong>Consistency</strong>: check your work for consistency: in spelling, in terminology, in typesetting.</li>
  <li><strong>Acronyms</strong>: before submitting, carefully check if every acronym has been introduced (and also if it is only introduced the very first time the term is mentioned).</li>
</ul>

<h2 id="writing-an-introduction-for-a-research-paper">Writing an Introduction for a Research Paper</h2>

<ul>
  <li>Motivation: Give a real-world motivation of a potential application of your work, or state which problem it solves. Then, give a high-level overview of the research area including several citations.</li>
  <li>Gap: What is the problem with existing approaches? Also, cite the 2-3 most closely related papers and state how you will overcome their limitations.</li>
  <li>Method: Give a brief high-level overview of the most important features of your method (or new dataset, if any).</li>
  <li>Contributions: A concise list of contributions (perhaps integrated with a Plan of the Paper, i.e., references to sections) should be at the end of your Introduction section. This part may be typeset as a bullet point list or as a list within the regular paragraph numbered with (1), (2), … or (a), (b), … or (i), (ii), … etc. – Pay special attention to the contributions, they are the currency in which your work is evaluated!</li>
</ul>

<h2 id="writing-tipps-for-bachelormaster-thesis">Writing Tipps for Bachelor/Master thesis</h2>

<p>The typical reader of your thesis is a peer: imagine a Bachelor/Master student who has taken the same courses as you have. The thesis should describe everything on a level so they can understand.</p>

<p><strong>Template and Length:</strong> Please use the template provided in <a href="https://git.rz.uni-augsburg.de/coling-a/coling-ba-ma-template">our uni-internal GitLab</a>.
There is no strict length limit, as the ideal length of a thesis strongly depends on your topic. Typically, a Bachelor thesis should not contain less than 30 pages (of text, excluding figures) and no more than 50 pages, while a Master thesis should contain between 40 and 60 pages. Keep in mind that writing <em>concisely</em> is much harder than producing large amounts of text, so your grade does not depend on the number of pages you produce. In fact, if the writing is repetitive, this will be evaluated negatively!</p>

<p><strong>Structure:</strong></p>

<p>How to write the <strong>Introduction</strong> for your thesis:</p>
<ul>
  <li>Start with writing the background, method, and experiments chapters. You will very likely have to edit the Introduction chapter less this way.</li>
  <li>Assume that if a reader has only limited time, they will only read the Introduction. What should they know about your work? (They should know about the most important points!)</li>
  <li>Start with the description and motivation of the problem your work is attempting to solve: This could either be a real-world problem or another technical issue for which your method is a prerequisite. (1-2 pages)</li>
  <li>Give a high-level introduction to your topic area (just what is necessary to understand the next point). (1-2 pages)</li>
  <li>Give a high-level description of the solution you will present in the thesis (dataset, method, findings). (2-3 pages)</li>
  <li>Give a plan of the paper (in the form of a list): what chapters will follow and what do they contain? (3-4 sentences per chapter)</li>
  <li>Very important: give a list of contributions. Typically, a thesis will have 2-3 main contributions, focus on these, and explain them in about 10 sentences each. Which gap do they solve? What is novel about them? (about 1 page)</li>
</ul>

<p>How to write the <strong>Background</strong> chapter:</p>
<ul>
  <li>Your thesis should not simply reproduce textbook-level knowledge. Only the general background that is important to your thesis should be explained.</li>
  <li>1-2 sub-chapters ideally explain the most closely related specific prior work in detail. This will usually be around 5-10 papers for a Bachelor thesis, and 10-20 papers for a Master thesis.</li>
</ul>

<p>How to write the <strong>Dataset/Method</strong> chapter:</p>
<ul>
  <li>Describe every step you took like a “recipe” (not a temporal report). Your work should be reproducible.</li>
  <li>For datasets, include corpus statistics, distribution of classes, etc., in the form of a table or diagrams.</li>
  <li>When including diagrams and preparing them with another program: make sure to embed PDFs.</li>
  <li>Talk to us how to write this best for your specific topic.</li>
</ul>

<p>How to write the <strong>Experimental Results</strong> chapter:</p>
<ul>
  <li>Typically, our work is empirical, testing a hypothesis on some data. Clearly state what hypothesis you are testing in each experiment.</li>
  <li>Describe the setup, hyperparameter tuning, which model checkpoint or prompt you used, your datasplits, …</li>
  <li>Include a few meaningful baselines for comparison: always predicting the majority class, etc.</li>
  <li>Report significance tests or standard variation of your results.</li>
  <li>Discuss each result: what does the finding mean?</li>
</ul>

<p>How to write the <strong>Discussion</strong> chapter:</p>
<ul>
  <li>Discuss your results and findings from a bird’s eye view: What do they implicate? What went wrong? What was unexpected?</li>
  <li>You can also integrate an <strong>Outlook</strong> section here (the title becomes <strong>Discussion and Outlook</strong>): What would be the next steps? What could be improved?</li>
</ul>

<p>How to write the <strong>Conclusion</strong> chapter:</p>
<ul>
  <li>This can be brief (1-2) pages, just summarize your main findings.</li>
</ul>

<p>What does into the <strong>Appendix</strong>:</p>
<ul>
  <li>long examples or prompts (short version in Figure in main text!)</li>
  <li>long pseudocode</li>
  <li>annotation guidelines</li>
  <li>huge tables with detailed experimental results (short version in main text!)</li>
  <li>… ask if in doubt! …</li>
</ul>

<h2 id="citation-style">Citation style</h2>

<p>Please use the ACL Natbib citation format in Latex, which you can find <a href="https://github.com/acl-org/acl-style-files">here</a>.
Most relevant papers can be found in the <a href="https://aclanthology.org/">ACL Anthology</a>, please use the bibtex entries from this source.
If a paper is not in the ACL Anthology, try finding a high-quality Bibtex entry in <a href="https://dblp.org/">dlpb</a> - also for ArXiv papers!
If an ArXiv paper has a follow-up version accepted for publication at a conference or journal, you should cite this peer-reviewed version.
Only if you cannot find an entry in these sources, resort to other sources for Bibtex entries (e.g., Google Scholar).
Please pay special attention to the instructions of how to integrate citations into the text. Always refer to the authors, <em>they</em> state/propose/found something, not “the paper.” This means you should write: “Friedrich et al. (2020) found that …” (use the <code class="language-plaintext highlighter-rouge">\citet</code> command!) rather than: “In (Friedrich et al., 2020) it was found that …” If a citation refers to a full sentence, put it <em>within</em> the sentence, e.g., “XYZ can be achieved using really small language models (Friedrich et al., 2031).”
There also exists a <code class="language-plaintext highlighter-rouge">\citeauthor</code> command that will output only the name of the author, e.g., “Friedrich et al.”. You can use this if you have already cited the paper WITH the year in a nearby context. In fact, in this case it is recommended rather than repeating the full citation. However, be careful: it is bad style to keep citing the same paper over and over again. Look carefully: are there other papers you could/should cite as well in the context? Can you use other formulations that still make it clear which work you are referring to without mentioning the name of the authors over and over again?</p>

<h2 id="cross-references">Cross-References</h2>

<p>Your readers should be able to read your thesis or papers in linear order. (As a child, I used to like the kind of crime stories where, as a reader, I could play detective by jumping to particular pages: a thesis or paper should not fall into this genre!)
Use cross-references with care. Forward references are almost always bad: as a rule of thumb, only use them if absolutely necessary (this should not be the case more than 2 times per thesis/article).
Instead, when you feel inclined to refer forwards, describe the important points in 2-3 sentences such that your reader will be able to follow the ideas you will describe next.
Backward references are okay, but only use them if absolutely necessary as well.</p>

<h2 id="examples-and-abbreviations">Examples and abbreviations</h2>

<p>I tend to just put commas before and after every <em>e.g.</em> and <em>i.e.</em> (following American English spelling).
For a good explanation, see <a href="https://en.wiktionary.org/wiki/i.e.">here</a> and <a href="https://en.wiktionary.org/wiki/e.g.">here</a>.</p>

<h2 id="typesetting">Typesetting</h2>

<p>There is no unique wrong or right way for typesetting. Here are my recommendations.</p>

<ul>
  <li>Introducing acronyms: “In this paper, we leverage Large Language Models (LLMs) to …”</li>
  <li>IMPORTANT! Introduce acronyms the <strong>first</strong> time you use them (in the main text, avoid acronyms in abstract if possible) and then use them <em>consistently</em>!</li>
  <li><em>italic</em>: use for introducing (highly) specific terminology, e.g., “This work deals with distinguishing <em>discourse type</em> vs. the <em>text type</em>.” Use italics only the first time you mention the term.</li>
  <li><strong>bold</strong>: use for paragraph headings when using the <code class="language-plaintext highlighter-rouge">\paragraph{...}</code> command. Sometimes, another phrase in the first sentence of a paragraph may function as a paragraph heading. Occasionally, it may make sense to boldface this phrase. Use this with care!</li>
  <li>Tables, figures etc. should be placed at the top of pages that do not have a chapter heading. Captions go below the table/figure, etc.</li>
  <li>For diagrams, if you are not familiar with <a href="https://tikz.net/">TikZ</a>, we recommend <a href="https://app.diagrams.net/">draw.io</a> (embed your diagrams as PDF).</li>
  <li>Use chapters (sections, subsections, paragraphs). In your text, you can write “In this chapter, …” or “In this section, …”, do not write “In this paragraph” or “In this subsection.” Preferred: title case for chapter and section headings, you can use lower case spelling for subsections and lower.</li>
  <li>“quotes”: In scientific text on NLP or CL, you will occasionally cite text examples. I tend to put them into quotes (no matter whether they are full sentences or just words or phrases). Use the <code class="language-plaintext highlighter-rouge">csquotes</code> package in Latex and use the <code class="language-plaintext highlighter-rouge">\enquote{...}</code> command for beautiful quotes and easy typesetting.</li>
</ul>

<p>For longer examples, I use the following Latex environments (I do not typeset these numbered linguistic examples in parentheses).</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>% Put this in your pre-amble
\usepackage{enumitem}

\newcounter{examplecounter}
\newlength\myLeftmargin
\setlength{\myLeftmargin}{0pt}
\newenvironment{example}
{\refstepcounter{examplecounter}
	\begin{samepage}
		\begin{indentation}{0.2em}{0em} % first: indent on left side, second: indent on right side
			\setlength{\parsep}{0pt}
			\setlength{\parskip}{0pt}
			\setlength{\itemsep}{0pt}
			\begin{list}{(\arabic{examplecounter})}%
				{\global\addtolength\myLeftmargin{\parindent}%
					\global\addtolength\myLeftmargin{\parindent}
					\setlength\leftmargin{\myLeftmargin}%
				}
				\item\relax}
			{\end{list}
		\end{indentation}
	\end{samepage}
}

\newcommand{\labex}[1]{\label{ex:#1}}
\newcommand{\eref}[1]{(\ref{ex:#1})}

% inside the document:
\begin{example} \small
On Monday, NASA {\bf announced} that signs of liquid water {\bf have been found} on Mars. The Mars Reconnaissance Orbiter spacecraft {\bf found} evidence of the liquid on the Martian surface, [...].\\% in long dark spots on the Red Planet thought to be formed because of water flow. \\
(\itshape https://en.wikinews.org/wiki/NASA\_announces\_water\_on\_Mars)
%In a news conference, NASA's planetary science director, Jim Green {\bf said}, "We now know Mars was once a planet very much like Earth with warm salty seas and fresh water lakes [...] but something has happened to Mars, it lost its water."
\label{ex:report}
\end{example}

As illustrated by \eref{report}, ...

</code></pre></div></div>

<h2 id="writing-assistance">Writing assistance</h2>

<p>NLP develops writing assistance tools and spellcheckers - it would hence be nonsensical not to make use of them!
However, be aware that you need to double-check their output: is it really what you intended? Is nothing hallucinated?
If you use spellcheckers like <a href="https://www.grammarly.com/grammar-check">Grammarly</a> that also help you to improve your style, you do not need to mention this.
If you make use of generative AI like ChatGPT for writing, please add a section in your appendix where you describe which sections you have composed with their help, and also reflect your experience with the writing tool (1-2 pages minimum).</p>

<h2 id="references">References</h2>

<p>Here are some additional references that may be helpful.</p>

<p>An <a href="https://blog.apastyle.org/apastyle/2011/08/punctuating-around-quotation-marks.html">overview</a> of where to put quotation marks in English.</p>

<p>A frequent mistake in English writing is to use non-parallel constructions in the same sentence. A good explanation can be found <a href="https://www.lingualbox.com/blog/parallel-structure-a-key-to-effective-english-writing">here</a>. Checking this out is highly recommended!</p>

<p><a href="http://users.ece.cmu.edu/~koopman/essays/abstract.html">This page</a> provides a nice summary of how to write good abstracts for your research papers.</p>

<p>I highly recommend <a href="https://www.worldscientific.com/worldscibooks/10.1142/q0232">this book</a> on science research writing for non-native speakers of English as a must-read for PhD students in our field.</p>

<p>Super useful <a href="https://www.phrasebank.manchester.ac.uk/">academic phrasebank</a> (thanks to Hanna Schmück for the suggestion).</p>

<h2 id="tense-connectives-and-other-important-writing-issues">Tense, connectives, and other important writing issues</h2>

<p><a href="http://annefried.github.io/files/Scientific_Writing_Exercises_09-22.pdf">Scientific Writing Cheat Sheet + Exercises (PDF)</a>, developed for PhD students, but certainly also useful for Bachelor/Master students.</p>

<p>As a rule of thumb, in scientific writing in the NLP domain, use <strong>present tense</strong>. There are only a few exceptions: when reporting on a process that you performed only once (e.g., data collection or annotation), you can use past tense, and (mostly in the conclusion) when you talk about achievements, use present perfect (e.g., “In this thesis, we have demonstrated that …”).</p>

<p>Make use of the <a href="https://en.wikipedia.org/wiki/Serial_comma">Oxford comma</a>.</p>

<h2 id="acknowledgment">Acknowledgment</h2>

<p>I learned many of these scientific writing concepts and tricks from my absolutely marvelous PhD and scientific writing mentor <a href="https://www.colorado.edu/linguistics/alexis-palmer">Alexis Palmer</a>.</p>]]></content><author><name>Annemarie Friedrich</name></author><category term="nlp" /><category term="scientific writing" /><summary type="html"><![CDATA[This blog post summarizes some important hints and tricks when writing a scientific paper or a thesis in the area of natural language processing (NLP) or computational linguistics (CL). Of course, they are only recommendations and the ideal solution depends on your concrete topic.]]></summary></entry><entry><title type="html">AIhub Blog Post on Methods to Address Class Imbalance in Deep-Learning Based NLP</title><link href="https://annefried.github.io/posts/2023-03-30-class_imbalance" rel="alternate" type="text/html" title="AIhub Blog Post on Methods to Address Class Imbalance in Deep-Learning Based NLP" /><published>2023-03-30T00:00:00-07:00</published><updated>2023-03-30T00:00:00-07:00</updated><id>https://annefried.github.io/posts/class_imbalance</id><content type="html" xml:base="https://annefried.github.io/posts/2023-03-30-class_imbalance"><![CDATA[<p>Sophie Henning (PhD student at the Bosch Center for Artificial Intelligence (BCAI)) and I got invited to write a blogpost for AIhub about our EACL 2023 paper on methods for addressing class imbalance in DL-based NLP. If you ever encountered the problem of your classifier only paying attention to the majority classes and wondered what to do about this, this article is for you. :)</p>

<p><a href="https://aihub.org/2023/03/30/methods-for-addressing-class-imbalance-in-deep-learning-based-natural-language-processing/">AIhub BlogPost</a></p>]]></content><author><name>Annemarie Friedrich</name></author><category term="deep learning" /><category term="class imbalance" /><summary type="html"><![CDATA[Sophie Henning (PhD student at the Bosch Center for Artificial Intelligence (BCAI)) and I got invited to write a blogpost for AIhub about our EACL 2023 paper on methods for addressing class imbalance in DL-based NLP. If you ever encountered the problem of your classifier only paying attention to the majority classes and wondered what to do about this, this article is for you. :)]]></summary></entry><entry><title type="html">Some thoughts on Area Chairing / Meta-Reviewing</title><link href="https://annefried.github.io/posts/2020-12-01-area_chairing" rel="alternate" type="text/html" title="Some thoughts on Area Chairing / Meta-Reviewing" /><published>2020-12-01T00:00:00-08:00</published><updated>2020-12-01T00:00:00-08:00</updated><id>https://annefried.github.io/posts/area_chairing</id><content type="html" xml:base="https://annefried.github.io/posts/2020-12-01-area_chairing"><![CDATA[<p><em>Thanks to Alexis Palmer, Casey Kennington and Nathan Schneider for their comments on an earlier version of this.</em></p>

<h2 id="preamble-why-i-am-writing-this">PREAMBLE: Why I Am Writing This.</h2>
<p>This year, I was asked for the first time to be an Area Chair both for ACL and EMNLP. It was a great insight and experience. First of all, I would like to state that I am totally convinced that the vast majority of ACs takes this job very seriously and does a great job. However, in my opinion, there is room for improvement in the instructions to ACs, which would be especially helpful for first-timers like myself this year. Luckily, I was working with really awesome Senior ACs in each case who were very responsive, encouraging and helpful during my process in creating my own recipe for how to do a good area chairing job. Here, I’d like to share my experience. Maybe it will be helpful for someone in the future.</p>

<h2 id="what-are-the-responsibilities-of-an-area-chair--meta-reviewer">What are the Responsibilities of an Area Chair / Meta-Reviewer?</h2>
<p>In the ACL community, this role is responsible for looking through the initial reviews (usually three), making a recommendation whether the paper is ready to be published, and writing a short so-called meta-review summarizing the major points of the reviews as well as the reviewers’ discussion.</p>

<h2 id="point-1-the-timing">POINT 1: The Timing.</h2>
<p>It’s a great honor to be asked to be an AC. It’s great for your CV. It gives you lots of great insights. If you’re asked, it means you have reached some kind of senior level in you research area. So, if you’re asked, by all means, support the community and say yes. IF you can make the time. Because it really is time-intensive, even if you know the subject well. The tasks were, with a realistic assumption how long it takes (at least the first time):</p>

<p>Checking potential problems in reviewer assignment, around 10 minutes per paper. I had almost 20 papers one time, so this took me around 2 hours. No problem this far.</p>

<p>As the reviews and author responses come in, you are supposed to start a discussion. This means actually reading all reviews thoroughly and starting a discussion among reviewers, at least in those cases where you’d like to have further opinions or want reviewers to elaborate / discuss their points. Count minimum half an hour per paper on average, which means … block an entire day in your calendar, just to be safe.</p>

<p>Soon afterwards, you have to write the meta-reviews. Luckily, some papers had been withdrawn at this point. Still, 14 meta-reviews to write. Of course, you don’t have to provide an advisor’s view on each not-yet-ready paper. However, behind many, if not most papers, there are PhD students who spent a lot of time on their papers. The least these young people entering our community deserve is a polite and useful feedback on why their paper was not accepted/needs an other iteration of research or writing. Remember, they might review your work one day. More thoughts on how to achieve this below. The point of this list was … right, timing. Above, I called myself a new senior-level person in my research area, however, to be honest, there are still many things that I need to look up. Sometimes, for example, I know that some previous paper did something but does this really mean that the paper under review is not novel? And so on. Plus, in a few cases, you won’t be able to rely on the initial reviews in that case. If you are already more senior than me, you might not have this problem. Well, I wanted to do a good job so I looked up some things here and there. Just to be clear, in one case I ended up recommending a paper whose merit the initial reviewers did not see, but to convince the SAC and the PCs, I actually did dig a bit deeper and almost wrote a review myself. I think it is just the case that these days many reviewers are relatively new to the field and not all initial reviews are to be counted upon. But hey, this is why we are there, and we shouldn’t just say “back luck with the reviewers,” but instead be the safety net in that case. Also, this doesn’t mean you have to read all 20 papers, but in that example I started to wonder from just reading the abstract and the reviews. So, if a case is a clear accept or reject, you might be done in an hour. But for some cases, I really spent two hours or more. In sum, I think I spent around 4-5 days writing meta-reviews, including my entire weekend. If you’re more  senior than me, you might be faster, but reading the abstract, reviews, author response, discussion, and writing the review takes at least 45 minutes per paper. (Apparently Nathan is faster. :))</p>

<p>Two take-aways: (1) If you agree to be an AC, block some days in your calendar early. (2) Conference organizers (note: ARR may have alleviated some things here), please allow more time for thorough meta-reviewing in the future, e.g., two weeks would be helpful.</p>

<p>From the SAC point of view, it is extremely helpful if you communicate in a timely fashion. Even if it’s to say “I haven’t gotten to this yet but I will tomorrow.” When coordinating amongst 15 ACs (for example) it would be tedious if half of them didn’t respond to an email and needed reminding. (Thanks to Nathan Schneider for this point.)</p>

<h2 id="point-2-the-recipe-for-writing-a-meta-review">POINT 2. The Recipe for Writing a Meta-Review.</h2>
<p>As an author, it is important to me to notice that the AC has actually gone through the relevant material and arrives at a recommendation for the right reasons. Any decision above your level will most likely happen just based on your scores and text. (Note: Do not explicitly state your recommendation in the text, though, because the PC decision may differ from yours.) This is my suggested structure for meta-reviews. Of course, some types of papers may need a slightly different structure and I diverged from my scheme if I didn’t have to say anything important for some categories. It’s more like a template to be modified.</p>

<p><strong>SUMMARY</strong>. 2-3 sentences (but no more) summarizing what the paper is about, as usual.</p>

<p><strong>STRENGTHS and CONTRIBUTIONS</strong>. Here, I am stating the reasons to accept the paper. I am not repeating the contributions list from the paper, instead, I am just noting the points that I would consider real contributions.</p>

<p><strong>WEAKNESSES</strong>. Things to be improved, especially for rejected papers. For instance, a methodological flaw or a missed comparison with particular prior work. (Keep in mind what has already been published at the submission deadline of your conference, though.)</p>

<p><strong>WRITING</strong>. Some comments on whether the writing is clear enough or needs improvement.</p>

<p><strong>REVIEWS and DISCUSSION</strong>. Summarize what the reviewers recommend, especially what happened in the discussion. Make sure to keep this double-blind. Possibly also add in information that was received via the author response.</p>

<p><strong>SUMMARY</strong>. A summary of your evaluation of the paper, e.g., “needs only minor modifications” or “would profit from another iteration.” As stated above, do not make your recommendation explicit here, though.</p>

<p><strong>ADDITIONAL COMMENTS</strong>. In cases where you disagree with the initial reviews, especially when recommending rejection or when agreeing only with some of the reviewers, state some points that make clear that you took a closer look at the work.</p>

<p>ACs have more time than SACs to evaluate the reviews and papers. Hence, do not to be afraid to take sides in weighing the pros and cons for papers with mixed reviews. It is super helpful for SACs if the meta-review has comments like “one reviewer asked for X, but in my view this is beyond the scope of the paper” or “reviewers complained about some points that were unclear but they are minor details and should be easy to clarify for the camera-ready.” (Again thanks to Nathan Schneider for this point.)</p>

<h2 id="point-3-final-notes">POINT 3. Final Notes.</h2>
<p>Thanks for reading up to here. If you made it here, I am sure you will be a great AC because you are interested in this subject. Please let me know what you think about these ideas. Finally, I hope you will have many great insights and interactions with the community while AC’ing (meta-reviewing)!</p>]]></content><author><name>Annemarie Friedrich</name></author><category term="area chairing" /><category term="acl" /><summary type="html"><![CDATA[Thanks to Alexis Palmer, Casey Kennington and Nathan Schneider for their comments on an earlier version of this.]]></summary></entry><entry><title type="html">Scientific Presentations</title><link href="https://annefried.github.io/posts/2020-10-05-scientific_presentations" rel="alternate" type="text/html" title="Scientific Presentations" /><published>2020-10-05T00:00:00-07:00</published><updated>2020-10-05T00:00:00-07:00</updated><id>https://annefried.github.io/posts/scientific_presentations</id><content type="html" xml:base="https://annefried.github.io/posts/2020-10-05-scientific_presentations"><![CDATA[<p>My slides on <a href="(http://annefried.github.io/files/scientific-presentations_friedrich.pdf) (admittedly using a slightly outdated beamer template, but the major points still hold :)">how to give effective scientific presentations</a>.</p>

<p>Some suggestions for <a href="http://annefried.github.io/files/FeedbackRules_Suggestion.pdf">feedback rules</a> for dry runs etc.</p>]]></content><author><name>Annemarie Friedrich</name></author><category term="scientific presentations" /><category term="nlp" /><summary type="html"><![CDATA[My slides on how to give effective scientific presentations.]]></summary></entry></feed>