Articles by convexstrictly
10

Retire the Abstractions (stanford.edu)

7

Gigatoken: Fastest Tokenizer (twitter.com/marcelroed)

4

Gemini Flash 2.0 Thinking Experimental (github.com/googleapis)

1

What Questions Are in the Chinese College Entrance Exam? (cherylwu.substack.com)

1

Generative AI Is Not Going to Build Your Engineering Team for You (stackoverflow.blog)

2

Building GPT2o – Part 1: Audio (medium.com/nivibilla)

3

The Geometry of Categorical and Hierarchical Concepts in Large Language Models (arxiv.org)

3

OpenAI says it has begun training a new flagship A.I. model (nytimes.com)

19

California residents: call your legislators about AI bill SB 1047 (twitter.com/chrislengerich)

3

LISA: Layerwise Importance Sampling for Memory-Efficient LLM Fine-Tuning (arxiv.org)

2

NTIA AI Open Model Weights RFC (regulations.gov)

1

Mechanics of Next Token Prediction with Self-Attention (arxiv.org)

1

Dive Deeper into Yi-9B (huggingface.co)

3

You can now train a 70B language model at home (answer.ai)

1

Shape Suffixes – Good Coding Style (medium.com/noamshazeer)

3

Star Trek prompt optimal for grade school math on Llama-70B (twitter.com/emollick)

1

(US Dept of Commerce) NTIA Solicits Comments on Open-Weight AI Models (commerce.gov)

4

BitDelta: Your Fine-Tune May Only Be Worth One Bit (arxiv.org)

37

Time is encoded in the weights of finetuned language models (arxiv.org)

2

Zoology 1: Measuring and Improving Recall in Efficient Language Models (stanford.edu)

2

TinyGSM: Achieving >80% on GSM8k with small language models (arxiv.org)

2

Androids built to meet the labor demands (1x.tech)

6

Sam Altman will likely start another company with researchers leaving OpenAI (twitter.com/emilychangtv)

492

Three senior researchers have resigned from OpenAI

1

Ron Conway disapproves of Sam Altman's firing (twitter.com/ronconway)

111

Sutskever: OpenAI board doing its mission to build AGI that benefits all (twitter.com/garymarcus)

3

Kara Swisher: OpenAI dev day and store were "pushing too fast (twitter.com/karaswisher)

1

GPT4 coding regression claims misleading (twitter.com/si_boehm)

3

Model 4 bit inference 4.2x faster than 16 bit with full HF support (twitter.com/tim_dettmers)

3

SqueezeLLM: Dense-and-Sparse Quantization (arxiv.org)

2

Inference-Time Intervention: Eliciting Truthful Answers from a Language Model (arxiv.org)

2

Orca: Progressive Learning from Complex Explanation Traces of GPT-4 (arxiv.org)

4

SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression (arxiv.org)

2

Azure GPT 3.5 completion endpoint bumps HumanEval from <50% to 74% (twitter.com/amanrsanger)

76

Falcon 40B LLM (which beats Llama) now Apache 2.0 (twitter.com/thom_wolf)

3

Tim Dettmers: QLoRA finetunes a 65B model on a single 48 GB GPU (twitter.com/tim_dettmers)