Researchers have developed a framework to calculate rigorous bounds on the probability of large language models generating harmful output in response to a given prompt. This approach utilizes Clopper-Pearson confidence intervals to obtain probably approximately correct bounds, providing a novel application of statistical methods to mitigate risks associated with language models. The proposed algorithm exploits features in the latent space to improve bound accuracy. By establishing a probabilistic safety bound, developers can better assess and manage the potential risks of large language models, which is crucial given the security implications that trail the rapid development of these models1. This matters to practitioners as it enables them to make more informed decisions about model deployment and risk mitigation, ultimately shaping the capability and risk surfaces of large language models.