<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Joseph Bejjani · Blog</title><description>Writing by Joseph Bejjani on AI alignment, reinforcement learning, and language models.</description><link>https://josephbejjani.com/</link><item><title>Can AI Write New War &amp; Peace? A ‘Novel’ Without End in a World of LLMs</title><link>https://drive.google.com/file/d/1YLREG4HTe2M99jvLGHPrFv9z-y90MWSG/view?usp=sharing</link><guid isPermaLink="true">https://drive.google.com/file/d/1YLREG4HTe2M99jvLGHPrFv9z-y90MWSG/view?usp=sharing</guid><description>A final paper for Slavic 118: prompting an LLM to write new chapters of War and Peace at two junctures in the novel, and examining how its chapters and self-reflections differ from Tolstoy.</description><pubDate>Sun, 10 May 2026 00:00:00 GMT</pubDate></item><item><title>Inoculating Language Models Against Misalignment</title><link>https://josephbejjani.com/blog/misalignment-inoculation/</link><guid isPermaLink="true">https://josephbejjani.com/blog/misalignment-inoculation/</guid><description>Inoculation prompting can mitigate emergent misalignment but may also create backdoor triggers.</description><pubDate>Fri, 30 Jan 2026 00:00:00 GMT</pubDate></item><item><title>Do LLMs understand their adversarial prompts?</title><link>https://josephbejjani.com/blog/perplexing-poem-prompt/</link><guid isPermaLink="true">https://josephbejjani.com/blog/perplexing-poem-prompt/</guid><description>Discovering perplexing prompts that generate poems—then asking the LLM to explain.</description><pubDate>Mon, 22 Dec 2025 00:00:00 GMT</pubDate></item><item><title>When Agents Prefer Hacking To Failure: Evaluating Misalignment Under Pressure</title><link>https://www.lesswrong.com/posts/AJANBeJb2p39su6F9/cs2881r-week-8-when-agents-prefer-hacking-to-failure</link><guid isPermaLink="true">https://www.lesswrong.com/posts/AJANBeJb2p39su6F9/cs2881r-week-8-when-agents-prefer-hacking-to-failure</guid><description>What do agents do when they face obstacles to a goal? If the only path to a goal requires misaligned action, will they choose it or accept failure? We build off Anthropic&apos;s work on Agentic Misalignment to investigate these questions in an agentic coding environment.</description><pubDate>Sun, 09 Nov 2025 00:00:00 GMT</pubDate></item><item><title>What I&apos;ve learned doing RL with JAX</title><link>https://josephbejjani.com/blog/mechagogue-jax/</link><guid isPermaLink="true">https://josephbejjani.com/blog/mechagogue-jax/</guid><description>Some of my experiences while working on mechagogue, a reinforcement learning repository with from-scratch JAX implementations of classic RL algorithms.</description><pubDate>Wed, 25 Jun 2025 00:00:00 GMT</pubDate></item></channel></rss>