Question Solved

What should I include in a statement of purpose for a Fulbright application?

Back to Forum
Starting 15 pts 0 followers
Yenepoya (Deemed to be University) · Posted

Background

I am a second-year PhD student working on multilingual NLP at NITK. My research involves building named entity recognition systems for code-mixed Telugu-English text, but I am struggling with the fine-tuning process for large language models.

The problem

I have around 8,000 annotated sentences, which I know is relatively small for fine-tuning something like XLM-R. When I fine-tune with default settings, I get reasonable results on validation but the model overfits badly by epoch 5.

What I have tried

  • Reduced learning rate to 1e-5
  • Added dropout at 0.3
  • Early stopping with patience=3
  • Tried LoRA fine-tuning as an alternative

Has anyone dealt with similar low-resource scenarios? Is LoRA actually recommended for NER tasks, or is full fine-tuning with heavy regularization the better path?

Any references to papers or GitHub repos with similar setups would be very helpful.

Sign in to join the discussion.

7 Replies

0
Jaya Lakshmanan Growing 50 pts · Accepted answer

I have a slightly different take from my experience in industry research. The reproducibility crisis is real but unevenly distributed. Fields with strong engineering culture (computational biology, ML with benchmarks) have actually improved significantly in the last 5 years. The bigger problem is in fields where data sharing is structurally difficult.clinical medicine, behavioral economics.

The ML community's move toward open code and reproducibility checklists has been genuinely effective.

0
Ravi Patel Starting 25 pts · Accepted answer

Seconding the recommendation for Zotero. Game changer for managing references across multiple projects.

0
Amit Joshi Starting 15 pts · Accepted answer

This question comes up a lot. The answer really depends on your specific field and what your committee values.

0
Replying to Amit Joshi
Pooja Nambiar Growing 115 pts · Accepted answer

Thank you! The reference to MuRIL is particularly useful.I had not considered it as an alternative to XLM-R for Indic languages.

0
Replying to Amit Joshi
Asha Pillai Distinguished 880 pts · Accepted answer

Thank you for the honest take. This is the kind of answer I was looking for.not the sanitized version.

0
Replying to Amit Joshi
Chirag Mehta Starting 15 pts · Accepted answer

I am in a very similar situation. Would you be willing to share the outline of your PMRF proposal? Not the content.just the section headings and approximate word allocation.

0
Dinesh Kulkarni Distinguished 730 pts · Accepted answer

+1 to everything said above. My experience was identical.