1
00:00:00,100 --> 00:00:06,820
So Claude Code shipped version 2.1.284,
and the default Sonnet changed underneath

2
00:00:06,880 --> 00:00:07,300
everybody.

3
00:00:09,140 --> 00:00:10,080
Um, no big banner.

4
00:00:10,660 --> 00:00:11,940
Just a line in the release notes.

5
00:00:12,200 --> 00:00:14,920
Like somebody swapping the engine while
you're parked at the servo.

6
00:00:15,760 --> 00:00:16,700
Which model are we talking?

7
00:00:16,750 --> 00:00:18,110
Sonnet 5.5.

8
00:00:18,310 --> 00:00:21,710
The model ID is claude sonnet five five.

9
00:00:21,850 --> 00:00:26,670
The release notes say it's now the default
Sonnet on the Anthropic API,

10
00:00:26,702 --> 00:00:29,150
with a one million token context window.

11
00:00:29,330 --> 00:00:34,910
Two dollars per million tokens in, ten
dollars out, and cache reads at twenty cents

12
00:00:34,970 --> 00:00:35,710
per million.

13
00:00:35,750 --> 00:00:39,670
And this quick breakdown is brought to you
by Jellypod AI, just so that's on the

14
00:00:39,727 --> 00:00:40,150
record.

15
00:00:40,270 --> 00:00:43,590
Right, twenty cents versus two dollars.

16
00:00:43,670 --> 00:00:45,270
That's a tenth, yeah?

17
00:00:45,572 --> 00:00:47,972
Yeah, a ninety percent discount on cached
input.

18
00:00:48,572 --> 00:00:50,432
And here's the bit I, I, I didn't expect.

19
00:00:50,952 --> 00:00:56,592
The migration guide says the minimum
cacheable prompt drops from 1,024 tokens down

20
00:00:56,612 --> 00:00:57,632
to 512.

21
00:00:57,708 --> 00:00:58,188
Hang on.

22
00:00:58,368 --> 00:01:00,028
Why does that matter for me?

23
00:01:00,215 --> 00:01:01,468
My prompts are big.

24
00:01:01,500 --> 00:01:02,500
Are they though?

25
00:01:02,573 --> 00:01:07,500
Think about a tight CLAUDE.md file, a few
terse project rules,

26
00:01:07,553 --> 00:01:09,820
or a subagent with a short system prompt.

27
00:01:09,913 --> 00:01:13,020
Those could sit under 1,024 tokens.

28
00:01:13,110 --> 00:01:17,820
Nothing in the sources says this happens
in Claude Code, so that's my inference.

29
00:01:17,900 --> 00:01:22,300
But a prompt that size couldn't be cached
before and now, in principle,

30
00:01:22,353 --> 00:01:22,940
it can.

31
00:01:23,218 --> 00:01:23,658
Ohh.

32
00:01:24,998 --> 00:01:30,378
So short loops, the ones where you ask,
tweak, ask again, they're the ones that get

33
00:01:30,438 --> 00:01:30,778
cheaper.

34
00:01:31,858 --> 00:01:33,558
Small stuff finally counts.

35
00:01:33,708 --> 00:01:34,028
Right.

36
00:01:34,048 --> 00:01:37,628
Now the runtime side, because this is
where old code might trip.

37
00:01:37,768 --> 00:01:41,308
Sonnet 5.5 runs adaptive thinking by
default.

38
00:01:41,436 --> 00:01:44,748
Send a request with no thinking field, and
it thinks.

39
00:01:44,750 --> 00:01:45,670
Fair enough.

40
00:01:45,758 --> 00:01:47,150
What if I don't want it to?

41
00:01:47,167 --> 00:01:49,967
There's a lowest setting called between
tools.

42
00:01:50,067 --> 00:01:53,887
The old disabled setting now returns a 400
error.

43
00:01:54,087 --> 00:01:55,647
And, uh, there's a twist.

44
00:01:55,780 --> 00:02:00,207
Notes longer than a sentence or two that
the model writes between tool calls come

45
00:02:00,255 --> 00:02:03,247
back inside thinking blocks, empty by
default.

46
00:02:03,434 --> 00:02:06,287
So an interface that shows those notes can
go quiet.

47
00:02:06,492 --> 00:02:08,872
Wait, so it's talking, I just can't hear
it.

48
00:02:08,958 --> 00:02:10,558
Pretty much.

49
00:02:10,583 --> 00:02:12,223
What about forcing a tool call?

50
00:02:12,363 --> 00:02:13,463
I've done that plenty.

51
00:02:13,500 --> 00:02:13,900
Gone.

52
00:02:14,040 --> 00:02:17,820
Forcing a tool with tool choice any, or
naming a specific tool,

53
00:02:17,900 --> 00:02:19,260
gets a 400 error.

54
00:02:19,400 --> 00:02:25,020
The guide says send auto and mark the tool
strict, so the input matches your schema.

55
00:02:25,132 --> 00:02:27,580
Then tell the model in the prompt when to
use it.

56
00:02:27,583 --> 00:02:31,583
Hmm, so less handcuffs, more, uh, a really
clear brief.

57
00:02:31,583 --> 00:02:32,383
Yeah, that's it.

58
00:02:32,403 --> 00:02:37,823
And on the Claude Code side, 2.1.284 adds
three keybinding actions.

59
00:02:37,926 --> 00:02:43,023
effortSlider:decreaseEffort,
increaseEffort, and toggleUltracode.

60
00:02:43,123 --> 00:02:47,823
You can rebind the arrows and Tab on the
slash effort slider in keybindings.json.

61
00:02:47,833 --> 00:02:49,273
And the levels are...?

62
00:02:49,292 --> 00:02:49,852
Five of them.

63
00:02:50,032 --> 00:02:53,452
Low, medium, high, xhigh, and max.

64
00:02:53,612 --> 00:02:58,812
The guide says the default on the API is
high, and for agentic coding it suggests

65
00:02:58,848 --> 00:03:01,292
starting at medium for well specified
tasks.

66
00:03:01,292 --> 00:03:02,892
Okay, so what's the catch?

67
00:03:02,959 --> 00:03:04,252
There's always a catch.

68
00:03:04,396 --> 00:03:08,332
Last time I skipped the fine print I, I, I
nearly took down a client's site.

69
00:03:08,493 --> 00:03:09,033
Tokens.

70
00:03:10,053 --> 00:03:13,053
Sonnet 5.5 uses Sonnet 5's tokenizer.

71
00:03:13,733 --> 00:03:20,493
Against Sonnet 4.6, Sonnet 4.5, and Haiku
4.5, the same text produces

72
00:03:20,553 --> 00:03:23,613
about thirty percent more tokens,
depending on the content.

73
00:03:24,473 --> 00:03:27,133
And thinking tokens are billed as output
tokens.

74
00:03:27,250 --> 00:03:30,050
So the ten dollar side of the price sheet
is where the, the,

75
00:03:30,070 --> 00:03:31,730
the surprise bill lives.

76
00:03:31,750 --> 00:03:32,150
Could be.

77
00:03:32,310 --> 00:03:35,670
The guide says recount your tokens and
re-baseline cost.

78
00:03:36,448 --> 00:03:36,668
Fair.

79
00:03:37,708 --> 00:03:39,208
Right, the good stuff in the release
notes.

80
00:03:39,628 --> 00:03:40,368
First one I like.

81
00:03:40,888 --> 00:03:45,548
Auto mode now has a Yes, but ask again
next time answer, when Claude wants to read

82
00:03:45,608 --> 00:03:47,068
outside your working directories.

83
00:03:47,167 --> 00:03:50,607
So you allow that one read, and it still
asks about the next.

84
00:03:50,625 --> 00:03:52,225
Exactly the middle ground I wanted.

85
00:03:52,417 --> 00:03:55,985
Then there's slash mcp reconnect all.

86
00:03:56,105 --> 00:04:00,865
One command to retry every MCP server that
failed to connect or needs

87
00:04:00,924 --> 00:04:01,825
authentication.

88
00:04:01,833 --> 00:04:03,673
No more poking at them one by one.

89
00:04:03,793 --> 00:04:06,873
And the compaction fix, that one's sneaky
good.

90
00:04:06,976 --> 00:04:11,033
Before, Prompt is too long errors could
stick around even after compacting.

91
00:04:11,042 --> 00:04:12,122
Been there.

92
00:04:12,197 --> 00:04:14,402
Compact, and it's still too fat.

93
00:04:14,417 --> 00:04:19,137
Now if the compacted request is still too
long, Claude Code compacts once more,

94
00:04:19,167 --> 00:04:20,977
keeping less of the recent conversation.

95
00:04:21,140 --> 00:04:25,720
So the upgrade is a cheaper cache, a
chattier brain, and a bigger token bill to

96
00:04:25,760 --> 00:04:26,060
watch.

97
00:04:26,920 --> 00:04:27,720
I can live with that.

