Signal
Advertise
Signal
Advertise
Sign in
Submit
Discover trends that matter
Trending repositories
Daily
Weekly
Monthly
Yearly
Live mentions
Topics
GitHub trending
Repositories
Developers
Insights
Stats
Log in
avbiswas/finetuning_recipes — GitHub trending stats & insights | Trendshift
Bifrost
Omnigraph
VModal Mobile SDK
avbiswas/finetuning_recipes
#
NLP
Visit GitHub
Like avbiswas/finetuning_recipes, 0 likes
0
Bookmark avbiswas/finetuning_recipes, 0 bookmarks
0
Python
186
27
1 contributors
Social mentions
Recent discussions about this repository across the web
Deep Learning bros and sisters, I literally have a video series that teaches how to train an SLM through CPT -> SFT -> DPO -> RL Also covers synthetic data gen and writing harnesses for custom tasks,…
@neural_avb · x.com
Recently trained a tiny 135M LM through - Continued Pretraining (CPT) - Supervised Finetuning (SFT) - Preference Optimization (DPO) - Reinforcement Learning (GRPO) Training data gen, Unsloth, HF,…
@neural_avb · x.com
Recently trained a tiny 135M SLM through CPT -> SFT -> DPO -> RL for a youtube series. Recommend all ML devs/creators to work on something like it coz this journey taught me so much! The full…
@neural_avb · x.com
Recently trained a tiny lil (135M) SLM through CPT -> SFT -> DPO -> RL for a youtube series DeepSeek-V4-Flash costs pennies and was responsible for 80% of the synthetic data, reasoning traces, and…
@neural_avb · x.com
My best new habit is to get my agent to document all the hacks and cheat-code I am using to train a model. I have logs for every hyperparam change, dataset upgrade, and what it resulted in. A very…
@neural_avb · x.com
Built this Tiny 135M Reasoning SLM step-by-step through CPT, SFT, DPO, and now RL. Put days into this data/model pipeline. And it fucking works! Performs narrow targeted tasks on research text at 300…
@neural_avb · x.com
Watch this 35 min visual guide on post-training Tiny Language Models. How to prepare preference datasets, finetune LMs to pick better trajectories, and evaluate their diversity + quality. Training…
@neural_avb · x.com
Releasing a super light Answer-eq Reward Model... You can RL train SLMs using this on QA tasks where verifiers are hard/expensive to design! Simple API: Inputs a candidate text, a GT reference, and…
@neural_avb · x.com
The paper-instructions dataset now comes with a subset of reasoning traces This is an awesome training dataset, curated with deepseek-v4-flash and qwen3.6-35B-A3B using text-albumentations. Costed me…
@neural_avb · x.com
So cool that this papers dataset got ~270 downloads last month on HF. It has surpassed ~1000 lifetime. I am working on v2 with more high quality data, as well as more diverse tasks. I will also be…
@neural_avb · x.com
Load more
Repository activities
repository's daily and monthly activities across stars, forks, merged PRs, issues, and closed issues