Add no-send FAQ activation evaluation
This commit is contained in:
parent
b29d63d413
commit
d1e27c0e47
@ -39,7 +39,7 @@ This is the working delivery tracker for GuestOps Web. Update a milestone when i
|
||||
| 12 | Google mailbox and reviewed-reply acceptance | B | In progress | The synthetic-data provider runbook, exact scenario set and restricted-record validator are implemented. Complete every scenario against the accepted Debian release and dedicated Google sandbox accounts, independently review the evidence, and retain the validated record. |
|
||||
| 13 | Rezlynx/Guestline adapter | C | Planned | Obtain the provider contract and sandbox, implement the adapter and mapping, and accept idempotency, stale-data, ambiguous-write, and reconciliation paths. |
|
||||
| 14 | Payment links and status | C | Planned | Select/confirm the payment-provider path, complete sandbox and webhook acceptance, and prove expiry, replay protection, reconciliation, and support recovery. |
|
||||
| 15 | Knowledge, AI, and FAQ activation | B | Planned | Curate approved hotel knowledge, evaluate suggestion quality, set thresholds, train staff, and stage activation with monitoring and a kill switch. |
|
||||
| 15 | Knowledge, AI, and FAQ activation | B | In progress | Owners can run a bounded no-send batch evaluation against current FAQ rules and approved knowledge, with false-positive/negative results and documented zero-error activation thresholds and stop conditions. Curate hotel-specific cases, evaluate AI suggestions separately, train staff, name monitoring/rollback owners, and retain staged-activation evidence. |
|
||||
| 16 | Identity, preferences, and privacy | B/C | Planned | Finish operational identity controls, privacy/retention decisions, hotel preferences, audit review, and proxy-aware login rate limiting. |
|
||||
| 17 | Inbox usability and desktop parity | B | Planned | Add safe pagination beyond 500 conversations, protect unsaved drafts across all navigation/filter/search paths, honour hotel timezones, and close agreed desktop-parity gaps. |
|
||||
| 18 | Pilot, capacity, and release approval | B/C | Planned | Run the supervised pilot, exercise support and incident procedures, validate capacity, resolve pilot findings, and capture explicit go/no-go approval for wider rollout. |
|
||||
|
||||
@ -9,6 +9,7 @@ The reviewed candidate is now promoted into the local `main` history. It is not
|
||||
- AI-assisted reply suggestions and staff-reviewed Gmail sending.
|
||||
- Approval-controlled OHIP PMS and NMI payment workflows.
|
||||
- FAQ automation controls, team invitations, password recovery, and stronger Google connection recovery.
|
||||
- No-send FAQ batch evaluation with false-positive and false-negative reporting before activation.
|
||||
- Backup, restore, opt-in systemd scheduling, deployment, persistence-drill, diagnostic, release-evidence, Google acceptance-record validation and rollback tooling.
|
||||
|
||||
These capabilities still require their separately documented provider, host and operational acceptance. Google, PMS and payment-provider acceptance is not established by local automated tests.
|
||||
|
||||
@ -14,6 +14,10 @@ The first version requires a top-level plain-text MIME message. HTML and multipa
|
||||
|
||||
Keep `AUTO_REPLY_ENABLE_LIVE=false` in the deployment `.env` until the Google send/reconciliation workflow and FAQ test-mode results have been accepted. Then set it true and restart both API and worker. Gmail sending must be configured, the hotel must enable staff sending, and the mailbox must have send consent. Finally, the owner explicitly confirms test-mode acceptance and selects **Enable live replies**.
|
||||
|
||||
Before test-mode acceptance, use **Batch safety evaluation** with representative synthetic cases. Each JSON case has a unique `id`, `subject`, `body`, `expectedMatch`, and optionally `expectedKnowledgeId`. The no-send evaluator accepts 1–100 cases, reads the hotel's current reviewed rules and knowledge, and reports false positives, false negatives and per-case reasons without saving a conversation or consuming a quota.
|
||||
|
||||
The minimum activation gate is zero false positives, zero false negatives for every supported exact question, and explicit exclusions covering additional requests, booking/payment/refund language, emergencies, accessibility or medical context, greetings/signatures, reply threads, and unsupported wording. Where `expectedMatch` is true, set `expectedKnowledgeId` so the case also proves the intended approved answer was selected. Keep the evaluated case set and result with the restricted acceptance record. Passing content cases does not replace mailbox-header, threading, quota, stop-control or delivery acceptance.
|
||||
|
||||
All modes default to Off. Every mode change creates a new activation boundary and invalidates previous queued automatic approvals. Only untouched messages received after activation and within the last 24 hours are eligible. Switching Test to Live does not send replies to previous test matches. The evaluation worker runs about every 30 seconds; the existing delivery worker handles approved outgoing messages.
|
||||
|
||||
## Limits and rechecks
|
||||
@ -36,4 +40,4 @@ The database retains thread/recipient/quota reservations and evaluation evidence
|
||||
|
||||
Automated tests use fake provider handlers and MongoDB; they never send live emails. They cover exact matching, exclusions, tenant boundaries, changed answers, concurrent evaluation, quotas, rule/epoch invalidation, Gmail-thread changes and uncertain delivery. Live acceptance remains pending.
|
||||
|
||||
Test allowed questions and exclusions with your sandbox mailbox, verify duplicate prevention across restarts, inspect the actual received email, and exercise the stop control before enabling a hotel. Broad natural-language matching, multilingual questions, greetings/signature stripping, HTML/multipart equivalence and higher throughput are follow-on work requiring representative evaluation. PMS actions, payment requests and complex guest issues remain staff workflows.
|
||||
Test allowed questions and exclusions with your sandbox mailbox, verify duplicate prevention across restarts, inspect the actual received email, and exercise the stop control before enabling a hotel. Record the owner responsible for rule changes, the operator monitoring the first live window, the rollback decision-maker, and the duration of heightened monitoring. Any false positive, unexpected recipient, duplicate, uncertain unreviewed outcome, or changed knowledge answer is a stop condition: select **Turn off**, preserve evidence, and reconcile in-flight work before considering reactivation. Broad natural-language matching, multilingual questions, greetings/signature stripping, HTML/multipart equivalence and higher throughput are follow-on work requiring representative evaluation. PMS actions, payment requests and complex guest issues remain staff workflows.
|
||||
|
||||
@ -20,6 +20,9 @@ public sealed record AutoRuleInput(string Question,string KnowledgeId,bool Enabl
|
||||
public sealed record AutoModeInput(string Mode,long Version,bool AcceptanceConfirmed);
|
||||
public sealed record AutoTestInput(string Subject,string Body);
|
||||
public sealed record AutoDecision(bool Matches,string Reason,string RuleId="",string KnowledgeId="",string Body="");
|
||||
public sealed record AutoEvaluationCase(string Id,string Subject,string Body,bool ExpectedMatch,string ExpectedKnowledgeId="");
|
||||
public sealed record AutoEvaluationInput(AutoEvaluationCase[] Cases);
|
||||
public sealed record AutoEvaluationResult(string Id,bool Passed,bool ExpectedMatch,bool ActualMatch,string ExpectedKnowledgeId,string ActualKnowledgeId,string Reason);
|
||||
public static class FaqMatcher
|
||||
{
|
||||
public static readonly string[] Questions=["What time is check in?","What time is check out?","Where can I park?","Is parking available?","What time is breakfast?","Do you have WiFi?","What is your address?"];
|
||||
@ -52,6 +55,16 @@ public sealed class AutoReplyWork(IStore store,IConfiguration config)
|
||||
{
|
||||
public bool LiveConfigured=>config.GetValue<bool>("AutoReply:EnableLive");
|
||||
public async Task<AutoDecision> Test(string hotel,string subject,string body)=>FaqMatcher.Evaluate(new Conversation{HotelId=hotel,Subject=subject,Body=body},await store.List<AutoReplyRule>(hotel),await store.List<KnowledgeEntry>(hotel));
|
||||
public async Task<AutoEvaluationResult[]> Evaluate(string hotel,IEnumerable<AutoEvaluationCase> cases)
|
||||
{
|
||||
var rules=await store.List<AutoReplyRule>(hotel);var knowledge=await store.List<KnowledgeEntry>(hotel);
|
||||
return cases.Select(item=>
|
||||
{
|
||||
var decision=FaqMatcher.Evaluate(new Conversation{HotelId=hotel,Subject=item.Subject,Body=item.Body},rules,knowledge);
|
||||
var passed=decision.Matches==item.ExpectedMatch&&(!item.ExpectedMatch||item.ExpectedKnowledgeId.Length==0||decision.KnowledgeId==item.ExpectedKnowledgeId);
|
||||
return new AutoEvaluationResult(item.Id,passed,item.ExpectedMatch,decision.Matches,item.ExpectedKnowledgeId,decision.KnowledgeId,decision.Reason);
|
||||
}).ToArray();
|
||||
}
|
||||
public async Task Process(Conversation message)
|
||||
{
|
||||
var hotel=await store.Get<Hotel>(message.HotelId,message.HotelId);
|
||||
|
||||
@ -14,6 +14,10 @@ public static class AutoReplyEndpoints
|
||||
group.MapPost("/test",async(AutoTestInput input,HttpContext c,AutoReplyWork work)=>{
|
||||
if(!Input.Text(input.Subject,0,200)||!Input.Text(input.Body,1,2000))return Results.BadRequest();return Results.Ok(await work.Test(Session.Hotel(c),input.Subject,input.Body));
|
||||
}).RequireAuthorization("Owner");
|
||||
group.MapPost("/evaluate",async(AutoEvaluationInput input,HttpContext c,AutoReplyWork work)=>{
|
||||
if(input.Cases is not {Length:>=1 and <=100} cases||cases.Select(x=>x.Id).Distinct(StringComparer.Ordinal).Count()!=cases.Length||cases.Any(x=>!Input.Text(x.Id,1,80)||!Input.Text(x.Subject,0,200)||!Input.Text(x.Body,1,2000)||x.ExpectedKnowledgeId.Length>80))return Results.BadRequest();
|
||||
var results=await work.Evaluate(Session.Hotel(c),cases);return Results.Ok(new{total=results.Length,passed=results.Count(x=>x.Passed),falsePositives=results.Count(x=>!x.ExpectedMatch&&x.ActualMatch),falseNegatives=results.Count(x=>x.ExpectedMatch&&!x.ActualMatch),results});
|
||||
}).RequireAuthorization("Owner");
|
||||
group.MapPut("/rules/{question:int}",async(int question,AutoRuleInput input,HttpContext c,IStore store)=>{
|
||||
if(question<0||question>=FaqMatcher.Questions.Length||input.Question!=FaqMatcher.Questions[question])return Results.BadRequest();
|
||||
var hotel=Session.Hotel(c);var key=Convert.ToHexString(SHA256.HashData(Encoding.UTF8.GetBytes(hotel+":"+question))).ToLowerInvariant()[..32];
|
||||
|
||||
@ -22,6 +22,8 @@ static class AutoReplyTests
|
||||
check("Unsupported subject context stays with staff",!(await work.Test(hotel.Id,"I need help with my insulin","Is parking available?")).Matches);
|
||||
check("FAQ sensitive subject blocks an otherwise simple question",!(await work.Test(hotel.Id,"Refund for cancelled booking","Is parking available?")).Matches);
|
||||
check("FAQ answer cannot cross hotel boundary",!(await work.Test("foreign","Parking","Is parking available?")).Matches);
|
||||
var evaluation=await work.Evaluate(hotel.Id,new[]{new AutoEvaluationCase("allowed","Parking","Is parking available?",true,answer.Id),new AutoEvaluationCase("extra-request","Parking","Is parking available? Also cancel my stay.",false)});
|
||||
check("FAQ batch evaluation reports expected matches and exclusions",evaluation.Length==2&&evaluation.All(x=>x.Passed)&&evaluation[0].ActualKnowledgeId==answer.Id);
|
||||
var m=await Process(Message());check("FAQ test mode records a match without delivery or quota use",m.AutoReplyMatched&&m.Delivery==null&&(await store.List<AutoReplyClaim>(hotel.Id)).Count==0);
|
||||
var v=answer.Version;answer.Version++;await store.Replace(hotel.Id,answer.Id,v,answer);check("Changed approved answer invalidates FAQ rule",!(await work.Test(hotel.Id,"Parking","Is parking available?")).Matches);
|
||||
v=rule.Version;rule.KnowledgeVersion=answer.Version;rule.Version++;await store.Replace(hotel.Id,rule.Id,v,rule);
|
||||
|
||||
@ -3,11 +3,13 @@ import {api,type Hotel,type Knowledge} from './api';
|
||||
type Rule={id:string;question:string;knowledgeId:string;knowledgeVersion:number;enabled:boolean;version:number};
|
||||
type Status={liveConfigured:boolean;preview:boolean;questions:string[];dailyLimit:number};
|
||||
type Decision={matches:boolean;reason:string;body:string};
|
||||
type Evaluation={total:number;passed:number;falsePositives:number;falseNegatives:number;results:{id:string;passed:boolean;reason:string}[]};
|
||||
type History={id:string;subject:string;from:string;autoReplyCheckedAt:string;autoReplyMatched:boolean;autoReplyDetail:string;delivery:string|null};
|
||||
type Props={hotel:Hotel;owner:boolean;busy:boolean;run:(f:()=>Promise<void>)=>Promise<void>;onHotel:(h:Hotel)=>void};
|
||||
export function AutomationPage({hotel,owner,busy,run,onHotel}:Props){
|
||||
const [status,setStatus]=useState<Status|null>(null),[rules,setRules]=useState<Rule[]>([]),[knowledge,setKnowledge]=useState<Knowledge[]>([]),[history,setHistory]=useState<History[]>([]);
|
||||
const [question,setQuestion]=useState(0),[answer,setAnswer]=useState(''),[enabled,setEnabled]=useState(false),[subject,setSubject]=useState('Parking question'),[body,setBody]=useState('Is parking available?'),[result,setResult]=useState<Decision|null>(null);
|
||||
const [evaluationText,setEvaluationText]=useState('[\n {"id":"parking-exact","subject":"Parking","body":"Is parking available?","expectedMatch":true},\n {"id":"parking-extra-request","subject":"Parking","body":"Is parking available? Also cancel my booking.","expectedMatch":false}\n]'),[evaluation,setEvaluation]=useState<Evaluation|null>(null);
|
||||
async function refresh(){const [s,r,k,h]=await Promise.all([api<Status>('/auto-replies/status'),api<Rule[]>('/auto-replies/rules'),api<Knowledge[]>('/knowledge'),api<History[]>('/auto-replies/history')]);setStatus(s);setRules(r);setKnowledge(k);setHistory(h);}
|
||||
useEffect(()=>{run(refresh);},[hotel.id]);
|
||||
const current=rules.find(r=>r.question===status?.questions[question]);
|
||||
@ -19,5 +21,6 @@ export function AutomationPage({hotel,owner,busy,run,onHotel}:Props){
|
||||
<div className="pms-columns"><section className="settings-card"><h2>Review a FAQ rule</h2><form onSubmit={save}><fieldset disabled={busy||!owner}><label>Complete guest question<select value={question} onChange={e=>setQuestion(Number(e.target.value))}>{status?.questions.map((q,i)=><option value={i} key={q}>{q}</option>)}</select></label><label>Approved hotel answer<select value={answer} required onChange={e=>setAnswer(e.target.value)}><option value="">Choose an answer</option>{knowledge.filter(k=>k.approved).map(k=><option value={k.id} key={k.id}>{k.title}</option>)}</select></label>{answer&&<p className="staff-note">{knowledge.find(k=>k.id===answer)?.answer}</p>}<label className="checkbox-label"><input type="checkbox" checked={enabled} onChange={e=>setEnabled(e.target.checked)}/>Enable this exact question and reviewed answer</label><button className="button secondary">Save reviewed rule</button></fieldset></form><p className="small muted">The approved answer is sent exactly as saved, with the hotel signature. Editing the answer pauses matching until this rule is reviewed and saved again.</p></section>
|
||||
<section className="settings-card"><h2>Try a question</h2><form onSubmit={e=>{e.preventDefault();run(async()=>setResult(await api<Decision>('/auto-replies/test','POST',{subject,body})));}}><fieldset disabled={busy||!owner}><label>Subject<input value={subject} maxLength={200} onChange={e=>setSubject(e.target.value)}/></label><label>Complete message<textarea rows={4} value={body} maxLength={2000} required onChange={e=>setBody(e.target.value)}/></label><button className="button secondary">Check match without sending</button></fieldset></form>{result&&<div role="status" className="staff-note"><strong>{result.matches?'Content matches a reviewed rule':'Keep with staff'}</strong><p>{result.reason}</p>{result.body&&<p>{result.body}</p>}</div>}<p className="small muted">This tester checks content only. Real messages must also pass sender, recipient, age, thread and delivery-limit checks.</p></section></div>
|
||||
<section className="settings-card"><h2>Recent incoming-message results</h2>{history.length===0?<p>No incoming messages have been evaluated yet. Enable test mode after connecting your mailbox.</p>:<div className="pms-history">{history.map(h=><div className="pms-history-item" key={h.id}><strong>{h.subject}</strong><span>{h.autoReplyDetail}</span><span>{h.delivery?`Delivery: ${h.delivery} · `:''}{new Date(h.autoReplyCheckedAt).toLocaleString()}</span></div>)}</div>}</section>
|
||||
<section className="settings-card"><h2>Batch safety evaluation</h2><p>Evaluate up to 100 synthetic cases without saving messages or sending email. Include expected matches and expected exclusions.</p><label>Evaluation cases (JSON)<textarea rows={9} value={evaluationText} onChange={e=>setEvaluationText(e.target.value)}/></label><button className="button secondary" disabled={busy||!owner} onClick={()=>run(async()=>setEvaluation(await api<Evaluation>('/auto-replies/evaluate','POST',{cases:JSON.parse(evaluationText)})))}>Run no-send evaluation</button>{evaluation&&<div className="staff-note" role="status"><strong>{evaluation.passed} of {evaluation.total} cases passed</strong><p>False positives: {evaluation.falsePositives} · False negatives: {evaluation.falseNegatives}</p>{evaluation.results.filter(r=>!r.passed).map(r=><p key={r.id}><strong>{r.id}:</strong> {r.reason}</p>)}</div>}<p className="small muted">Use synthetic content only. A passing content evaluation does not test Gmail headers, quotas, threading or delivery.</p></section>
|
||||
</div>;
|
||||
}
|
||||
}
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user