Reinforcement learning personalization

강화학습 개인화

Use observed rewards to adjust which message is tried next.

···
html
<div class="learn"><div class="headline">ROUND <b id="round">1</b> <span>OBSERVE → UPDATE</span></div><div class="arm"><label>HELP</label><div class="rail"><i id="helpbar"></i></div><b id="helpnum">50%</b></div><div class="arm"><label>OFFER</label><div class="rail"><i id="offerbar"></i></div><b id="offernum">50%</b></div><div class="reward" id="reward">REWARD: HELP +1</div></div>
css
.learn{width:min(91%,590px);font:800 clamp(10px,3.2vmin,15px)/1.2 var(--font-sans)}.headline{border-bottom:2px solid var(--fg);padding-bottom:clamp(8px,2vmin,16px);display:flex;gap:6px}.headline span{margin-left:auto;color:var(--muted)}.arm{display:grid;grid-template-columns:clamp(40px,11vmin,75px) 1fr 35px;align-items:center;gap:8px;margin:clamp(13px,4vmin,25px) 0}.rail{height:clamp(19px,5vmin,32px);background:var(--line)}.rail i{display:block;height:100%;width:50%;background:var(--accent);transition:width .5s}.arm:nth-of-type(3) .rail i{background:var(--fg)}.arm b{text-align:right}.reward{background:var(--surface);border:2px solid var(--line);padding:clamp(7px,2vmin,14px);text-align:center}
js
const rounds=[[1,50,50,'HELP +1'],[2,65,35,'HELP +1'],[3,58,42,'OFFER +1'],[4,72,28,'HELP +1']];let n=0;function show(){const [round,help,offer,reward]=rounds[n];document.querySelector('#round').textContent=round;document.querySelector('#helpbar').style.width=help+'%';document.querySelector('#offerbar').style.width=offer+'%';document.querySelector('#helpnum').textContent=help+'%';document.querySelector('#offernum').textContent=offer+'%';document.querySelector('#reward').textContent='REWARD: '+reward;n=(n+1)%rounds.length}show();setInterval(show,1800);

Reinforcement learning personalization uses observed customer response as a reward signal to adjust future messages or offers. Braze's 2025 decisioning announcement describes continual experimentation and adaptation from customer behavior.

The demo shifts the next round's allocation after observing wins for two offers. It focuses on updating a policy from outcomes, while AI decisioning can describe a choice among current predicted scores.

A click-only reward can favor unwanted messages. Include long-term value and opt-outs when defining success.

When to use

Use it when offer allocation should keep learning from customer responses.

Open as page ↗