1
00:00:00,340 --> 00:00:03,960
Um, so, Ethan, you know how automated
security scanners are,

2
00:00:04,019 --> 00:00:08,660
like, notorious for yelling at you about
things that don't actually matter for your

3
00:00:08,740 --> 00:00:09,480
specific setup?

4
00:00:10,053 --> 00:00:11,133
Oh, absolutely.

5
00:00:11,653 --> 00:00:17,053
The classic critical severity page at
three in the morning for a dev dependency that

6
00:00:17,133 --> 00:00:18,693
isn't even exposed to production.

7
00:00:19,170 --> 00:00:19,690
Right!

8
00:00:19,870 --> 00:00:20,730
Exactly.

9
00:00:21,290 --> 00:00:27,810
Well, OpenAI just dropped version 0.1.29
of their open source codex security

10
00:00:27,910 --> 00:00:28,330
package.

11
00:00:28,770 --> 00:00:33,210
And, um, before we dive into the technical
details, shoutout to Jellypod for

12
00:00:33,270 --> 00:00:34,330
supporting our daily show!

13
00:00:35,050 --> 00:00:40,470
But yeah, this release is basically all
about giving teams actual control over how

14
00:00:40,530 --> 00:00:43,370
security findings and automated patches
get evaluated.

15
00:00:43,950 --> 00:00:45,350
Control in what sense?

16
00:00:45,550 --> 00:00:46,590
Like, custom rules?

17
00:00:47,067 --> 00:00:47,487
Yeah!

18
00:00:47,967 --> 00:00:52,687
So, pull request 940 introduced custom
patch validation prompts.

19
00:00:53,267 --> 00:00:57,627
When codex security finds a vulnerability
and tries to auto generate a fix,

20
00:00:58,127 --> 00:01:01,227
you can now pass in your own custom
validation prompt.

21
00:01:01,787 --> 00:01:05,067
So instead of just letting the model guess
if a patch is acceptable,

22
00:01:05,607 --> 00:01:07,307
you give it explicit constraints.

23
00:01:07,625 --> 00:01:08,505
Ah, I see.

24
00:01:08,665 --> 00:01:14,425
So if your team has a policy like, say,
never introduce new external dependencies,

25
00:01:14,612 --> 00:01:19,545
or you must match a specific coding
pattern, you can enforce that directly in the

26
00:01:19,602 --> 00:01:22,425
prompt before it even generates the pull
request?

27
00:01:22,798 --> 00:01:23,618
Exactly.

28
00:01:23,918 --> 00:01:26,278
You guide the validation step directly.

29
00:01:26,738 --> 00:01:31,298
And then on the triage side, they added
custom severity classification rubrics.

30
00:01:31,667 --> 00:01:34,547
Wait, how does that work with existing
scans?

31
00:01:34,707 --> 00:01:36,787
Do you have to re run the whole scan?

32
00:01:37,232 --> 00:01:37,652
No!

33
00:01:37,992 --> 00:01:39,292
That is the cool part.

34
00:01:39,732 --> 00:01:45,172
You run npx at openai slash codex security
classify severity,

35
00:01:45,672 --> 00:01:49,272
pass it your scan ID, and point it to a
markdown rubric file,

36
00:01:49,672 --> 00:01:52,352
like slash path to policy dot md.

37
00:01:52,932 --> 00:01:58,092
It re assesses the finding severities
against your internal org policies without

38
00:01:58,192 --> 00:02:01,752
altering the sealed scan artifacts or
touching the raw findings.

39
00:02:02,225 --> 00:02:02,545
Hmm.

40
00:02:03,105 --> 00:02:04,085
That is pretty clean.

41
00:02:04,645 --> 00:02:09,525
From my product days, keeping the raw
audit record sealed while overlaying business

42
00:02:09,625 --> 00:02:10,745
logic on top is...

43
00:02:11,425 --> 00:02:13,005
that's huge for compliance.

44
00:02:13,465 --> 00:02:16,145
It completely eliminates triage noise.

45
00:02:16,705 --> 00:02:21,445
Back when I was doing software testing,
getting a raw model finding mapped directly

46
00:02:21,505 --> 00:02:23,865
to an emergency page was a nightmare.

47
00:02:24,445 --> 00:02:28,005
Now, if a finding doesn't match your
custom rubric's high threshold,

48
00:02:28,265 --> 00:02:29,765
it doesn't wake anyone up.

49
00:02:30,125 --> 00:02:34,605
So how does this actually look in a
developer's day to day CLI workflow?

50
00:02:35,465 --> 00:02:42,465
So, first you run npx at openai slash
codex security scan dot to scan your

51
00:02:42,525 --> 00:02:43,305
current directory.

52
00:02:43,985 --> 00:02:48,345
Then when you run the patch generation
command, you pass your custom patch

53
00:02:48,385 --> 00:02:49,305
validation prompt.

54
00:02:49,905 --> 00:02:53,205
And before exporting those findings to
Linear or GitHub issues,

55
00:02:53,685 --> 00:02:57,145
you run that classify severity step with
your markdown rubric.

56
00:02:57,658 --> 00:02:59,038
But wait, are there catch...

57
00:02:59,218 --> 00:03:00,378
are there caveats here?

58
00:03:01,038 --> 00:03:03,018
What happens if you don't supply a rubric?

59
00:03:03,423 --> 00:03:06,743
If no rubric is supplied, it's completely
opt in.

60
00:03:07,183 --> 00:03:11,463
It just falls back to the original scan
severity without making any extra model

61
00:03:11,523 --> 00:03:14,203
calls, so you don't burn extra API budget.

62
00:03:14,863 --> 00:03:18,343
But for custom validation prompts, you
have to be careful.

63
00:03:18,903 --> 00:03:22,183
If your prompt constraints are too
contradictory or strict,

64
00:03:22,663 --> 00:03:25,803
the model might just refuse to generate a
patch altogether.

65
00:03:26,125 --> 00:03:27,005
Makes sense.

66
00:03:27,133 --> 00:03:30,845
Give it an impossible task and it just
throws its hands up.

67
00:03:30,989 --> 00:03:36,045
Were there any other quality of life
updates in 0.1.29?

68
00:03:36,403 --> 00:03:37,963
A bunch of practical ones!

69
00:03:38,443 --> 00:03:43,443
Pull request 944 automatically includes
Terraform infrastructure files,

70
00:03:43,643 --> 00:03:46,643
dot tf files, in scan inventories now.

71
00:03:47,243 --> 00:03:53,523
And pull request 932 streams large saved
scan JSON logs, so your memory doesn't

72
00:03:53,643 --> 00:03:56,003
spike when scanning massive repositories.

73
00:03:56,333 --> 00:04:01,373
Oh, and they added sortable tables to the
findings dashboard preview in pull request

74
00:04:01,453 --> 00:04:02,573
958, right?

75
00:04:03,063 --> 00:04:03,623
Yes!

76
00:04:04,543 --> 00:04:07,683
Makes sorting through stored scan runs way
easier.

77
00:04:08,120 --> 00:04:12,080
You know, looking at this whole release,
it feels like we're moving past just

78
00:04:12,140 --> 00:04:14,620
trusting raw AI outputs for security.

79
00:04:15,400 --> 00:04:20,860
Giving teams policy driven validation
prompts bridges that gap between automated

80
00:04:21,020 --> 00:04:23,500
patch creation and actual code review.

81
00:04:23,942 --> 00:04:25,162
Yeah, exactly.

82
00:04:25,442 --> 00:04:28,042
It's about guardrails, not just
automation.

83
00:04:28,742 --> 00:04:31,902
Well, that is 0.1.29 in a nutshell!

84
00:04:32,442 --> 00:04:33,122
Good chatting, Ethan.

85
00:04:33,542 --> 00:04:34,582
Talk soon.

