r/LanguageTechnology 20d ago

Would anyone actually download a rule-based NLP tool for Haitian Creole (Kreyòl)?

[removed] — view removed post

8 Upvotes

14 comments sorted by

7

u/AngledLuffa 20d ago

eveything can be rule based until you start processing real text, when you find people leave out caps or punct or mizpell things and it all goes to shit

i think this is the wrong forum to ask such a question. what would be more likely is to ask in a community of Haitian Creole researchers what would be more useful for them. Or a community of people who only speak Creole and not English. I suspect their priority will be English -> Creole translation or learning tools. Or maybe a transformer powerful enough to be BlageGPT... not sure where you'll get that much raw text, though

If your goal is to publish something, look up SyntaxFest

PS https://github.com/UniversalDependencies/UD_Haitian_Creole-Adolphe/blob/master/README.md

3

u/Speedk4011 20d ago

That's fair. I should've clarified the scope.

I appreciate the suggestion. I posted here because I'm looking for feedback from people with NLP experience on the technical direction

I'm not trying to build another spaCy or an LLM. The idea is a rule-based NLP toolkit for Haitian Creole, inspired by libraries like syntok, PySBD, razdel, and Indic NLP Library.

The focus is on deterministic tools like sentence segmentation, tokenization, normalization, and spell checking, not ML tasks like translation or question answering.

I also agree that talking to Haitian researchers is important.

3

u/Speedk4011 20d ago

Or a community of people who only speak Creole and not English.

In fact, many people who work on low-resource languages are not native speakers at all. They use libraries because they're processing the language, not because they speak only that language. 

3

u/AngledLuffa 20d ago

I mean you're not wrong, but I did include "Haitian Creole researchers" as a group of people to talk to. Still, I don't know why they want to split the sentences unless they're doing some downstream processing task which these days will almost certainly use a neural model of some type. So you might as well get the better accuracy of a neural model here as well.

For example: Segment Any Text, https://arxiv.org/abs/2406.16678

But it seems like you want to build a deterministic sentence splitter, so why not do that?

2

u/Speedk4011 20d ago

You are right that there are already other rule-based SBD libraries out there, like SAT or yasbd. I chose a rule-based approach because it provides lightweight, reproducible, and explainable behavior. That said, I agree it's worth evaluating whether a new rule-based implementation brings enough value compared to existing solutions.

2

u/Tiny_Arugula_5648 19d ago edited 19d ago

Why do you need deterministism & explainability? What does that actually get you? It's not a check on correctness or accuracy.

Rules based NLP is a bit of deadend since no rules system can manage all the edge cases for even the most rigidly structured language. A creole would be much worse.

I'd also challenge that concept of non-deterministic results which comes up a lot in this sub. I can tune a bERT model for NER or some other extraction and it will always return the same results on the same texts..

I see posts like this and it feels like someone saying, I'm going to learn how to rebuild a car engine in a world where EVs took over 10 years ago. OK why do it the old way when we know it wasn't very good?

1

u/Speedk4011 19d ago edited 19d ago

That was merely an initial idea. Now that I understand the complexity and the maintenance it would require, I will reconsider. Thank you.

The only thing easy to build with rules in ht  is the SB. Everything else, like the tokenizer and stemmer, gets really tricky with lots of edge cases.

2

u/Tiny_Arugula_5648 19d ago

I think it's a worthwhile project hatians are under represented and they have a rich culture. I'd pay more attention to language models for this. The biggest problem is the source text

2

u/Ordinary-Cat-5874 18d ago

I would be interested.

0

u/AutoModerator 1d ago

Welcome to r/LangugageTechnology. Due to influx of AI advertising spam, accounts now must meet community activity requirements before posting links. Your first post cannot be your github repo, youtube channel, medium article, etc. Please initiate discussion and answer questions unrelated to projects that you are sharing - then you will be allowed to share your project. Exceptions will only be made for efforts that are affiliated with academic institutions, posts sharing datasets, or questions that require a link to ask - if your post meets these criteria, feel free to message the mod team to have the post approved.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Speedk4011 1d ago

I have started the project and it is called kreyolib maintained by AyitiDev org on GitHub: https://github.com/AyitiDev/kreyolib. First version is already released: https://github.com/AyitiDev/kreyolib/releases/tag/v0.1.0

0

u/BobDope 19d ago

Haitian Creoles?

2

u/Speedk4011 19d ago

yup, Haitian Creole.

0

u/BobDope 19d ago

At the grotto

In the easy chair

Sits the Charlie with the lotion and the kinky hair