Model techniques map
Techniques

specific method · not yet filed

defense-in-depth with layered input/output classifiers

source
1
model
1
labs adopt it
0
strongest
mentioned

How sources treat it

One count per evidence span, weakest treatment to strongest.

mentioned 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

are best addressed with defense-in-depth rather than relying on the model's refusals alone. Common downstream moderation tools, such as Llama Guard, are compatible with Inkling and can be layered around the model to catch jailbreak attempts, filter unsafe outputs, and enforce use-case-specific policies.

mentionedotherThinking Machines Lab