1
00:00:00,159 --> 00:00:05,679
You know, up until, uh, up until this
week, tweaking a custom skill or a plugin for

2
00:00:05,779 --> 00:00:10,820
Claude Code felt a bit like, er, pushing a
code patch at midnight without running

3
00:00:10,840 --> 00:00:11,679
tests first.

4
00:00:12,340 --> 00:00:16,319
You just, you just sort of crossed your
fingers and hoped your model wouldn't

5
00:00:16,340 --> 00:00:17,760
hallucinate its way off a cliff.

6
00:00:19,554 --> 00:00:22,074
Man, I have definitely felt that pain.

7
00:00:22,134 --> 00:00:26,435
You tweak one prompt line, and suddenly
your whole custom workflow stops returning

8
00:00:26,474 --> 00:00:29,594
proper JSON, or it drops a tool call
completely.

9
00:00:29,667 --> 00:00:30,067
Right!

10
00:00:30,147 --> 00:00:32,547
It was pure non deterministic chaos.

11
00:00:32,767 --> 00:00:37,427
But look, this episode is brought to you
by Jellypod AI, and we have got some

12
00:00:37,480 --> 00:00:42,867
properly good news from the latest release
of Claude Code, version 2.1.269.

13
00:00:42,875 --> 00:00:47,515
Yeah, 2.1.269 finally gives us a real
testing harness for plugins.

14
00:00:47,627 --> 00:00:51,915
Anthropic added a dedicated CLI command,
claude plugin eval.

15
00:00:51,917 --> 00:00:56,317
Wait, so how does, uh, how does claude
plugin eval actually work in practice?

16
00:00:56,477 --> 00:00:58,237
Is it like a full test runner?

17
00:00:58,833 --> 00:00:59,633
Pretty much!

18
00:00:59,753 --> 00:01:04,433
You run a plugin's eval suite against
Claude Code, and it runs automated benchmark

19
00:01:04,502 --> 00:01:05,073
passes.

20
00:01:05,185 --> 00:01:08,353
Then it emits scored, reproducible
outputs.

21
00:01:08,433 --> 00:01:13,633
You get both a raw JSON file with the
metrics and a formatted HTML report to view in

22
00:01:13,649 --> 00:01:14,353
your browser.

23
00:01:14,721 --> 00:01:17,561
Scored JSON and an HTML report.

24
00:01:18,801 --> 00:01:20,861
That is massive for indie builders.

25
00:01:21,361 --> 00:01:24,561
So instead of just guessing if your prompt
tweak made the tool better,

26
00:01:25,121 --> 00:01:26,321
you actually get a score?

27
00:01:26,542 --> 00:01:27,262
Exactly.

28
00:01:27,302 --> 00:01:30,782
You can catch regressions before pushing
changes to your team.

29
00:01:30,809 --> 00:01:35,902
If your plugin's score drops from 95
percent to 70 percent on tool accuracy,

30
00:01:35,942 --> 00:01:38,302
you see it instantly in the HTML report.

31
00:01:38,333 --> 00:01:39,773
That is clean as a whistle.

32
00:01:39,933 --> 00:01:44,333
But, er, I reckon there is a catch with
running automated evals over and over,

33
00:01:44,353 --> 00:01:44,733
isn't there?

34
00:01:44,946 --> 00:01:46,733
Token cost?

35
00:01:47,529 --> 00:01:48,509
Oh, absolutely.

36
00:01:49,089 --> 00:01:54,589
If you run a massive eval suite with fifty
test cases, every single test is making

37
00:01:54,609 --> 00:01:55,610
real model calls.

38
00:01:56,130 --> 00:01:58,769
That token burn adds up fast if you put it
on a tight loop.

39
00:01:58,992 --> 00:01:59,632
Yeah, right!

40
00:02:00,053 --> 00:02:03,452
You don't want to accidentally fry your
API quota while you are grabbing a white

41
00:02:03,592 --> 00:02:04,273
flat coffee.

42
00:02:04,813 --> 00:02:09,152
You have gotta keep an eye on your usage
limits when kicking off big benchmark runs.

43
00:02:09,250 --> 00:02:09,730
For sure.

44
00:02:09,790 --> 00:02:15,410
And speaking of running lots of things at
once, this 2.1.269 update added another

45
00:02:15,479 --> 00:02:17,330
really cool setting for workflows.

46
00:02:17,442 --> 00:02:20,850
Have you seen
CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS?

47
00:02:20,875 --> 00:02:22,315
Uh, no, actually.

48
00:02:22,331 --> 00:02:23,835
What does, what does that setting do?

49
00:02:23,875 --> 00:02:28,435
It lets you raise the Workflow tool's
concurrent agent limit anywhere from 1 all the

50
00:02:28,475 --> 00:02:32,915
way to 256 concurrent agents for inference
bound fan outs.

51
00:02:33,036 --> 00:02:34,277
Two hundred and fifty six?

52
00:02:35,276 --> 00:02:35,657
Wow.

53
00:02:35,936 --> 00:02:40,036
That is serious horsepower if you are
splitting up big jobs across a army of

54
00:02:40,156 --> 00:02:40,736
subagents.

55
00:02:40,917 --> 00:02:41,317
Right.

56
00:02:41,397 --> 00:02:46,437
If you have independent tasks like
scanning fifty repositories or running massive

57
00:02:46,490 --> 00:02:50,517
parallel checks, you can just unleash the
fan out without getting bottlenecked by

58
00:02:50,557 --> 00:02:51,637
the default runner cap.

59
00:02:51,667 --> 00:02:52,267
Crikey.

60
00:02:52,291 --> 00:02:53,427
That is proper fast.

61
00:02:53,607 --> 00:02:55,747
And what about the display side of things?

62
00:02:55,847 --> 00:02:58,547
Did they fix up the visual terminal output
at all?

63
00:02:58,583 --> 00:03:02,023
Yeah, they added /output style followed by
the style name.

64
00:03:02,123 --> 00:03:06,343
You can use it to list and switch prompt
layouts and output styles on the fly,

65
00:03:06,391 --> 00:03:09,143
even over Remote Control or in headless
cloud sessions.

66
00:03:09,167 --> 00:03:09,967
Ah, brilliant.

67
00:03:10,020 --> 00:03:14,607
So if you are running headless in CI or
controlling it remotely,

68
00:03:14,627 --> 00:03:17,647
you can strip down the formatting so it
doesn't clutter your logs.

69
00:03:17,667 --> 00:03:18,747
Precisely.

70
00:03:18,802 --> 00:03:24,467
Between reproducible plugin evaluation and
scaling up to 256 parallel agents,

71
00:03:24,513 --> 00:03:27,667
Claude Code is feeling way more like a
production engine now.

72
00:03:27,786 --> 00:03:31,626
Yeah, no more shipping untested code at
midnight and praying to the tech gods.

73
00:03:32,866 --> 00:03:36,446
Alright, that is 2.1.269 in a nutshell.

74
00:03:36,887 --> 00:03:37,627
Good chat, mate!

75
00:03:37,750 --> 00:03:39,030
Yeah, catch you next time!

