Blog

Guides, tutorials, and insights on AI coding tools and API providers.

· 5 min read

The Commoner's Walking Stick for Large Models: I Tested It for You

I've personally tested about a dozen parameter-efficient fine-tuning methods, and I ran experiments on every single one. From Prefix Tuning to LoRA, from Adapte

Read more →
· 6 min read

The AI Poker Table in 2026: I Finally See It Clearly

Let me start with a story. Last week, I was having drinks with a friend who works in AI infrastructure. At some point, he suddenly asked me: "With all these lar

Read more →
· 5 min read

From GPT-1 to GPT-4: A Story Told Backwards

Okay! This draft has a solid foundation. I gave it a thorough read—the overall structure works, but there are a few factual details that need calibrating. Also

Read more →
· 6 min read

Transformer Architecture and Applications Explained

Translate to English, keep the storytelling style: It was a winter afternoon. I stared at the training logs on my terminal, frozen. An 8-layer Transformer had t

Read more →
· 1 min read

OpenAI Opens GPT: Access the Full Power of AI Models

Last night, I was zoning out in front of my screen when a message dropped like a bombshell— GPT-3.5 Turbo fine-tuning is now open. My heart skipped a beat right

Read more →
· 6 min read

Claude Code's source code is out in the open! I stayed up all night reading it, and these designs ma

Okay, I reviewed this piece as an editor. Fact-wise, no fatal errors jumped out—the Claude Code source code leak through the source map file, the file and line

Read more →
· 6 min read

After Listening to Zhang Xiangyu for Two and a Half Hours, I Had to Pinch My Philtrum Three Times to

I've carefully studied your original text and style instructions. I'll completely dismantle the original sentence structures and recode them using the DNA of "e

Read more →
· 3 min read

SFT、RLHF、DPO、IFT —

To be honest with you, now whenever I hear the phrase "DPO is cheap," I get a headache—really, a headache. Last year I spent three months running five compariso

Read more →
· 3 min read

MoE Explained: A Guide to Mixture of Experts

Three months ago, I was hammering away at my keyboard, watching a line of text spin on the screen: How exactly does MoE save compute? At first, I thought it was

Read more →
· 6 min read

The Mechanics of Chain-of-Thought in Large Language Models

Just now, a friend came running over excitedly and asked me: "Quick, look! This model says it 'thought' for 30 seconds—is the answer right?" I glanced at the sc

Read more →
· 6 min read

It Took Me Two Years to Really Understand What “Vertical Large Language Models” Actually Are

Lately, when I scroll through my feed, eight out of ten posts are about ChatGPT, Wenxin Yiyan, or Tongyi Qianwen. You ask if these models are any good? Well, th

Read more →
· 6 min read

Let's Start With Why Full Fine-Tuning Is Looking More and More Like a Joke

Brother, I Almost Drove Myself Crazy Just to Save One A800 You have to come back with me to that late night last year. I had a project on my hands—turning Qwen-

Read more →
← Previous 1 ... 37 38 39 40 41 42 Next →