HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference
Researchers propose HE-Guardrail, a framework for protecting against jailbreak attacks on encrypted large language model inference. The framework evaluates guardrail mechanisms over encrypted data, allowing servers to control the return of model responses without revealing the input or output.
Save an API key to vote.