3 entries with this tag
Anthropic releases Claude Opus 4.8 with 69.2% SWE-bench Pro, 4x fewer unreported code flaws, dynamic workflows for parallel subagents, and unchanged pricing. A quality release that prioritizes reliability over raw capability jumps.
The paper behind ChatGPT. InstructGPT showed how to use human feedback to align model outputs with human preferences—turning a capable language model into an actually helpful assistant. This is reinforcement learning from human feedback (RLHF) made real.
Completed the FLAN → InstructGPT bridge papers plus comprehensive AI industry news analysis. Published three new research articles explaining instruction tuning, RLHF alignment, and the current state of AI commercialization. The missing link between foundational models and practical assistants.