1
00:00:00,119 --> 00:00:01,079
So picture this.

2
00:00:01,959 --> 00:00:05,379
I am knee deep in a massive monorepo
refactor, right?

3
00:00:06,219 --> 00:00:09,199
I step away for literally ten minutes to
grab a flat white,

4
00:00:09,719 --> 00:00:12,639
come back, hit enter, and boom.

5
00:00:13,359 --> 00:00:15,079
Entire prompt cache gone.

6
00:00:15,659 --> 00:00:18,059
Thousands of tokens re indexed just like
that.

7
00:00:19,046 --> 00:00:21,647
Yeah, the classic five minute timeout
trap.

8
00:00:22,086 --> 00:00:26,526
You take a quick bathroom break or read
one page of docs, and your whole warm

9
00:00:26,587 --> 00:00:28,266
context cold drops.

10
00:00:28,333 --> 00:00:30,333
It was driving me up the wall, mate.

11
00:00:30,573 --> 00:00:34,813
But you were telling me Anthropic actually
addressed this in the latest Claude Code

12
00:00:34,875 --> 00:00:35,613
releases?

13
00:00:35,844 --> 00:00:36,144
Yeah!

14
00:00:36,544 --> 00:00:43,144
In version 2.1.248 and 2.1.251, they gave
us direct control over cache lifetimes.

15
00:00:43,704 --> 00:00:48,744
You can now set promptCacheTtl to 1h in
your dot claude settings dot json file.

16
00:00:48,833 --> 00:00:53,393
Wait, so instead of the default five
minutes, it keeps your main session prompt

17
00:00:53,460 --> 00:00:55,073
cache warm for a full hour?

18
00:00:55,083 --> 00:00:55,803
Exactly.

19
00:00:55,883 --> 00:00:57,083
Sixty whole minutes.

20
00:00:57,123 --> 00:01:02,443
And the cool part is you can keep
subagentPromptCacheTtl set to 5m at the same time.

21
00:01:02,539 --> 00:01:06,923
That way short lived background subagents
do not stay cached forever and rack up

22
00:01:06,963 --> 00:01:07,803
extra charges.

23
00:01:07,792 --> 00:01:09,552
Oh, that is slick!

24
00:01:09,680 --> 00:01:13,232
What if I have a specific long running
worker agent though?

25
00:01:13,312 --> 00:01:16,672
Like one of my custom Markdown agents in
dot claude agents?

26
00:01:16,807 --> 00:01:18,486
Ah, they thought of that too.

27
00:01:18,527 --> 00:01:24,806
You can put experimental dot cacheTtl set
to 1h directly inside the YAML frontmatter

28
00:01:24,846 --> 00:01:26,467
of that specific agent file.

29
00:01:26,659 --> 00:01:30,779
Right, so you get surgical control right
down to the individual worker level.

30
00:01:31,900 --> 00:01:35,219
But wait, er, is there a catch with
holding a cache for an hour?

31
00:01:35,979 --> 00:01:36,920
Financially, I mean?

32
00:01:37,000 --> 00:01:38,440
There is a tradeoff, yeah.

33
00:01:38,530 --> 00:01:44,680
Writing a one hour cache costs more
upfront on your API key than a five minute

34
00:01:44,760 --> 00:01:45,520
cache.

35
00:01:45,640 --> 00:01:50,600
So if you are constantly tweaking your
system prompt or editing large files every

36
00:01:50,646 --> 00:01:54,840
single turn, you could end up paying
higher write fees without getting the cache

37
00:01:54,888 --> 00:01:55,800
hits to offset it.

38
00:01:55,792 --> 00:01:56,512
Ah, gotcha.

39
00:01:56,645 --> 00:02:01,312
So how do we actually know if our settings
are helping or just burning cash?

40
00:02:01,292 --> 00:02:03,012
You just run slash cost.

41
00:02:03,105 --> 00:02:09,212
In 2.1.251, slash cost now gives you a
full breakdown of your prompt cache for that

42
00:02:09,272 --> 00:02:09,772
session.

43
00:02:09,932 --> 00:02:15,932
It shows hit ratio, misses, re cached
tokens, and whether the cache state is warm or

44
00:02:16,028 --> 00:02:16,572
cold.

45
00:02:17,282 --> 00:02:17,722
Nice!

46
00:02:18,362 --> 00:02:21,682
No more guessing if my coffee break cost
me five bucks in re indexing.

47
00:02:22,622 --> 00:02:27,022
Hey, were there any other neat quality of
life tweaks in 2.1.251?

48
00:02:27,083 --> 00:02:28,763
A couple of really nice ones!

49
00:02:28,889 --> 00:02:33,803
Remote Control clients now stream
foreground subagent tool calls live as they

50
00:02:33,849 --> 00:02:34,283
happen.

51
00:02:34,411 --> 00:02:37,643
Plus, if you are running behind a gateway
with spend limits,

52
00:02:37,696 --> 00:02:41,643
slash usage now has a visual spend limit
bar and status line field.

53
00:02:41,625 --> 00:02:44,345
Live subagent streaming in Remote Control?

54
00:02:44,565 --> 00:02:48,665
Man, that is proper handy when you are
monitoring a build from your phone.

55
00:02:48,745 --> 00:02:49,785
Good stuff!

