Skip to content
View suhara's full-sized avatar

Block or report suhara

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Open-source library for scalable, reproducible evaluation of AI models and benchmarks.

Python 322 71 Updated Aug 7, 2026

Scalable toolkit for efficient model reinforcement

Python 1,891 502 Updated Aug 10, 2026

A dataset for training and evaluating LLMs on decision making about "when (not) to call" functions

Python 67 7 Updated Apr 29, 2025

🍷 Code for Noisy Pairing and Partial Supervision for Stylized Opinion Summarization (Iso et al; INLG 2024)

Python 3 Updated Aug 12, 2024

Python bindings for llama.cpp

Python 10,540 1,442 Updated Aug 9, 2026

We aim to provide the best references to search, select, and synthesize high-quality and large-quantity data for post-training your LLMs.

Python 66 5 Updated Oct 3, 2024

This repo contains the source code for RULER: What’s the Real Context Size of Your Long-Context Language Models?

Python 1,599 135 Updated Jul 22, 2026

日本語LLMまとめ - Overview of Japanese LLMs

TypeScript 1,424 45 Updated Aug 10, 2026

LLM training in simple, raw C/CUDA

Cuda 30,767 3,723 Updated Jun 26, 2025

Awesome LLM Papers and repos on very comprehensive topics.

223 24 Updated Aug 22, 2024

Scalable data pre processing and curation toolkit for LLMs

Python 1,706 313 Updated Aug 7, 2026
Python 13 3 Updated Feb 8, 2025

A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

Python 18,066 3,541 Updated Aug 9, 2026

Scalable toolkit for efficient model alignment

Python 853 109 Updated Oct 6, 2025

😈Awful AI is a curated list to track current scary usages of AI - hoping to raise awareness

7,529 263 Updated Feb 20, 2025

Python port of Moses tokenizer, truecaser and normalizer

Python 497 58 Updated Feb 6, 2026

A modular RL library to fine-tune language models to human preferences

Python 2,394 201 Updated Mar 1, 2024

Annotating Columns with Pre-trained Language Models

Python 35 13 Updated Jun 10, 2022

Repository to collect and categorize Grammatical Error Correction papers.

127 10 Updated Jan 30, 2026

A curated list of research papers and resources on Indonesian languages

41 3 Updated Mar 21, 2024

PyTorch code for "FactPEGASUS: Factuality-Aware Pre-training and Fine-tuning for Abstractive Summarization" (NAACL 2022)

Python 40 2 Updated Sep 15, 2022

xfspell — the Transformer Spell Checker

Shell 189 21 Updated Jun 18, 2020

This repository contains materials for the SIGIR 2022 tutorial on opinion summarization.

33 1 Updated Jul 20, 2022

🥥 Code & Data for Comparative Opinion Summarization via Collaborative Decoding (Iso et al; Findings of ACL 2022)

Python 23 3 Updated Mar 3, 2025

AutoPrompt: Automatic Prompt Construction for Masked Language Models.

Python 640 86 Updated Jul 17, 2026

This repository contains multiple notebooks (created using Google Colab) that transform data from a doccano format for use in training a Bi-LSTM-CRF, fine-tuning a transformer using custom labels, …

Jupyter Notebook 7 1 Updated Dec 8, 2021

This project uses Huggingface transformers GPT-2 to fine-tune text generation models based on lyric data to specific music genres.

Jupyter Notebook 7 1 Updated Dec 8, 2021

Judge a book by it's cover. Data from Open Library

Jupyter Notebook 1 Updated Dec 8, 2021

Final project for Topics in Computing

Jupyter Notebook 1 Updated Dec 8, 2021
Jupyter Notebook 1 Updated Dec 7, 2021
Next