- Santa Clara, CA
- https://yoshi-suhara.com
- @suhara
- in/yoshi-suhara
Stars
Open-source library for scalable, reproducible evaluation of AI models and benchmarks.
Scalable toolkit for efficient model reinforcement
A dataset for training and evaluating LLMs on decision making about "when (not) to call" functions
🍷 Code for Noisy Pairing and Partial Supervision for Stylized Opinion Summarization (Iso et al; INLG 2024)
We aim to provide the best references to search, select, and synthesize high-quality and large-quantity data for post-training your LLMs.
This repo contains the source code for RULER: What’s the Real Context Size of Your Long-Context Language Models?
日本語LLMまとめ - Overview of Japanese LLMs
Awesome LLM Papers and repos on very comprehensive topics.
Scalable data pre processing and curation toolkit for LLMs
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Scalable toolkit for efficient model alignment
😈Awful AI is a curated list to track current scary usages of AI - hoping to raise awareness
Python port of Moses tokenizer, truecaser and normalizer
A modular RL library to fine-tune language models to human preferences
Annotating Columns with Pre-trained Language Models
Repository to collect and categorize Grammatical Error Correction papers.
A curated list of research papers and resources on Indonesian languages
PyTorch code for "FactPEGASUS: Factuality-Aware Pre-training and Fine-tuning for Abstractive Summarization" (NAACL 2022)
This repository contains materials for the SIGIR 2022 tutorial on opinion summarization.
🥥 Code & Data for Comparative Opinion Summarization via Collaborative Decoding (Iso et al; Findings of ACL 2022)
AutoPrompt: Automatic Prompt Construction for Masked Language Models.
This repository contains multiple notebooks (created using Google Colab) that transform data from a doccano format for use in training a Bi-LSTM-CRF, fine-tuning a transformer using custom labels, …
This project uses Huggingface transformers GPT-2 to fine-tune text generation models based on lyric data to specific music genres.
Judge a book by it's cover. Data from Open Library
Final project for Topics in Computing