LSTM for the Win · Model results

Product review classification

A reproducible view of sentiment and product-topic classification on the newest controlled synthetic incoming batch. The dashboard reads the immutable run artifact generated by the model workflow.

RunLoading latest run…
Model executedLoading…
ScopeSynthetic · Incoming
Status Loading live data

Resolving the newest versioned workflow output from GitHub… View run files

Executive summary

Performance at a glance

Accuracy = correct predictions ÷ incoming reviews · 95% Wilson intervals shown below

Topic accuracy

Waiting for results

Both labels correct

Waiting for results

Incoming reviews

Unique incoming IDs

Benchmark interpretation

Loading the latest model comparison…

The interpretation is calculated from the canonical workflow run.

Model diagnostics

Where the models perform

Accuracy

Model comparison

Same incoming batch

Recall by expected label

Sentiment class performance

Higher is better

Error pattern

Sentiment confusion matrix

Rows: expected · Columns: predicted

Coverage

Topic volume and accuracy

Expected classes

Prediction behavior

Class balance and confidence

Sentiment counts

Expected vs. predicted

ExpectedPredicted

Average confidence

Certainty vs. accuracy

Confidence is the model probability assigned to its selected class. It is not the same as accuracy.

Review explorer

Inspect individual predictions

Loading reviews…

IDReviewSentimentTopicResult
Loading model results…
Page 1

How to read this dashboard

Data provenance and limitations

Data. The primary cards and explorer display the newest synthetic incoming batch. Runs retain an immutable synthetic benchmark for longitudinal comparison. New runs also include a separately sourced real-world sentiment benchmark.

Versioning. latest.json resolves the newest retained run. Each run exposes one immutable run.json; paper analyses and any later CSV or Parquet views are derived directly from that same run artifact.

Continual learning. A controlled fraction of incoming reviews marked goldtest is promoted after evaluation. The synthetic benchmark and external benchmark never enter training.

Scope. Training, incoming evaluation and the longitudinal benchmark remain synthetic. Independent external validation uses the CC BY 4.0 UCI Sentiment Labelled Sentences Amazon subset and applies to sentiment only. The external source has binary positive/negative labels and no compatible product-topic labels, so it does not establish external topic generalization or full three-class sentiment coverage.

Loading pipeline metadata…