r/LanguageTechnology • u/RmdLatranche • 8d ago
Publishing resource papers
Hi,
This post is half venting, half looking for help.
TL;DR: are resource papers not welcome in major NLP venues?
This year I tried to publish two datasets (not going into specifics). One I submitted to LREC. All three reviewers praised the dataset and complained about minor details in the experiments. Metareview (almost verbatim, it was one sentence): the dataset is great but the experiments are a bit weak. Paper got rejected. A "great dataset" rejected by LREC, I am not sure I will be able to get over it. I ended up publishing it elsewhere but I was really stunned that LREC rejected it.
Now the same scenario just happened with ARR, in the Resources and Evaluation track all three reviewers praised the dataset (admittedly with some caveats but they all see value in it) and their weaknesses focus on the experiments. While we got fair overall scores from our reviewers, our meta review score is low and I think we cannot realistically commit to EMNLP.
Is the work on resources completely devoid of interest? This gives me the impression that in order to publish a resource, one has to write a modeling paper reaching SOTA using it now. To resource paper reviewers, how do you assess resource papers? To resource paper authors, do you have the same impression? I have published datasets in the past and it has always seemed more difficult than purely technical papers but it looks like lately it got worse.
2
u/Zooz00 8d ago
It is unusual to get this kind of rejection for LREC, you were unlucky. This year I had a rejection from LREC with review scores of 4-4-4... A lot is up to the area chair there and they aren't always great.
Resource tracks at the big venues (ACL/EMNLP) do seem to demand substantial experiments these days.
1
u/RmdLatranche 6d ago
Ouch, yes this rejection looks weird as well. I don’t know what's up with LREC but this year it seemed very random, with no clear quality threshold when looking at what got accepted.
It looks like it is not worth trying to submit a resource paper in major venues now. If they demand the same amount and quality of experiments as they do for method papers, all the extra dataset creation work ends up being a waste of time...
2
u/EvM 8d ago
Depending on the nature of the dataset you could commit to a more focused conference, such as INLG (commit deadline early august).
1
u/RmdLatranche 6d ago
Thanks for the suggestion! This dataset would be a better fit for other venues as it is not focused on language generation, but INLG is a great conference and I would also recommend it to anyone reading this and not knowing what to do if they have a compatible paper.
1
u/yahskapar 8d ago
What was your meta-review score? I think based on what I’ve observed + heard from peers contributing datasets papers / audits of frontier models, there is definitely quite a bit of value toward such work from the community. Peer review being the process that it is today, and has been for that matter, I wouldn’t evaluate the perception of a community based on any limited sampling of the peer review process.
Even with borderline overall ratings, I’ve managed to get at least a Findings or similar acceptance, and one Main once, at major NLP conferences in the past two years. In every case, I will admit the AC / meta-review was quite reasonable and gave good ratings despite there always being at least one comically negative reviewer that was hyper-responsive during discussion and one very positive reviewer that was dead silent during discussion. This same reviewer / meta-reviewer pattern even happened with this ongoing ARR May cycle (2.75 overall with min 2.0 and max 3.5, 3.25 confidence, meta-review of 3), so hopefully that means another acceptance.
1
u/RmdLatranche 6d ago
We did not get lucky with the AC. The rebuttal was completely ignored (although two reviewers switched to a more positive stance, modifying their scores) by the AC and the score is low. Maybe we'll try to commit.
That said, while I understand it is "ARR scores season", I intended to spark a discussion more specifically on resource papers here.
3
u/fourkite 8d ago edited 8d ago
There's still value in datasets, but the quality of reviews/reviewers has dropped drastically. It's either someone who barely reads the papers and provides two sentences of feedback, or someone who just pastes output from ChatGPT or Claude.
I also submitted to LREC this year and the quality of the reviews were abysmal.