alli: how do we increase context windows for boyfriends? man is forgetful
Jeremy Nguyen:
- begin by allocating a clear role ("I just want to share, you don't have to solve the problem")
- use verbal xml tags ("oh, I gotta fill you in on this thing my boss said", "okay, that's the end of the boss thing")
- users think they want sycophancy, but they don't want it to be obvious
Shweta:
you don’t increase the context window, you schedule regular RLHF fine-tuning loops
Jeremy Nguyen:
- begin by allocating a clear role ("I just want to share, you don't have to solve the problem")
- use verbal xml tags ("oh, I gotta fill you in on this thing my boss said", "okay, that's the end of the boss thing")
- users think they want sycophancy, but they don't want it to be obvious
Shweta:
you don’t increase the context window, you schedule regular RLHF fine-tuning loops
X (formerly Twitter)
alli (@sonofalli) on X
how do we increase context windows for boyfriends? man is forgetful
❤9👍1😁1
りょう@おしり描きません:
#ガルクラ #リコリコ #girlsbandcry
いらっしゃいませ〜✨
喫茶ガルクラへようこそ〜♪
https://twitter.com/macqueen7777/status/1918578794610201047
#ガルクラ #リコリコ #girlsbandcry
いらっしゃいませ〜✨
喫茶ガルクラへようこそ〜♪
https://twitter.com/macqueen7777/status/1918578794610201047
🤣4🥰1
Forwarded from Hacker News
Time saved by AI offset by new work created, study suggests (Score: 151+ in 4 hours)
Link: https://readhacker.news/s/6tRYU
Comments: https://readhacker.news/c/6tRYU
Link: https://readhacker.news/s/6tRYU
Comments: https://readhacker.news/c/6tRYU
Ars Technica
Time saved by AI offset by new work created, study suggests
Survey of 2023–2024 data finds that AI created more tasks for 8.4 percent of workers.
👍1
Forwarded from Hacker News
AI code is legacy code? (Score: 150+ in 12 hours)
Link: https://readhacker.news/s/6tUYj
Comments: https://readhacker.news/c/6tUYj
Link: https://readhacker.news/s/6tUYj
Comments: https://readhacker.news/c/6tUYj
Text Incubation
AI code is legacy code from day one - Text Incubation
AI code is legacy code from day one 5/4/25 This was on the Hacker News front page at the time - I've included some of the more interesting comments in a section below. It seems like there are a few s…
Forwarded from 每日字体观察 / 每日字體觀察
「中文字体解秘组计划」第一阶段成果 - The Type
点此下载报告全文的 PDF 文件
2022 年,3type 发起了「中文字体解密组」的研究项目。研究项目的总体概念、研究方法和路径……经过三年时间,我们迎来了第一阶段的成果。
在「中文字体解密组」第一阶段,我们的研究对象是思源黑体简体中文版三个字重(ExtraLight,Regular 和 Heavy)中的 9169 个汉字字符形……我们招募了很多志愿者参与研究项目,手工为目标字符集中的汉字添加了各种不同的标签,并依据各种标签和字体文件本身包含的数据进行分析,试图解开集体无意识设计工作背后隐藏的规律。……
点此下载报告全文的 PDF 文件
❤1
Forwarded from Hacker News
Matrix-vector multiplication implemented in off-the-shelf DRAM for Low-Bit LLMs (Score: 151+ in 13 hours)
Link: https://readhacker.news/s/6tVHC
Comments: https://readhacker.news/c/6tVHC
Link: https://readhacker.news/s/6tVHC
Comments: https://readhacker.news/c/6tVHC
arXiv.org
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM...
General matrix-vector multiplication (GeMV) remains a critical latency bottleneck in large language model (LLM) inference, even with quantized low-bit models. Processing-Using-DRAM (PUD), an...
Hacker News
Matrix-vector multiplication implemented in off-the-shelf DRAM for Low-Bit LLMs (Score: 151+ in 13 hours) Link: https://readhacker.news/s/6tVHC Comments: https://readhacker.news/c/6tVHC
TLDR: 之前的研究发现通过不正确的指令时序,可以在市面上买到的 DRAM 模块内部执行并行度非常高的复制和位运算。这篇论文提出了提升数据存储和运算效率的方法,使得直接在内存中执行低精度大模型成为可能。研究发现这一做法相对 CPU 有两三倍的性能和能效提升。
🤯34🔥4🤔1
Forwarded from Twitter Picture Bot
🥰28