risan / sentiment-analysis
Fast, dependency-free sentiment analysis for PHP. VADER scoring for English and Indonesian.
Requires
- php: ^8.3
- ext-mbstring: *
Requires (Dev)
- laravel/pint: ^1.0
- pestphp/pest: ^4.0
- phpstan/phpstan: ^2.0
Suggests
None
Provides
None
Conflicts
None
Replaces
None
This package is auto-updated.
Last update: 2026-10-01 19:41:58 UTC
README
Scores English and Indonesian text as positive, negative or neutral. Pure PHP 8.3+, no API calls, no dependencies.
Documentation and live examples: https://sentiment-analysis.risanb.com
Install
composer require risan/sentiment-analysis
Requires PHP 8.3 or newer and the mbstring extension.
Usage
Analyze a text
Pass a string. You get a result with a label and a compound score.
use Risan\Sentiment\Sentiment; $result = Sentiment::analyze( 'This package is awesome!', ); $result->label; // Label::Positive $result->compound; // 0.6588
compound runs from -1 (very negative) to +1 (very positive). The label comes from compound.
Read the result
The three shares add up to about 1. Three methods check the label.
$result->positive; // 0.594 $result->neutral; // 0.406 $result->negative; // 0.0 $result->isPositive(); // true $result->isNegative(); // false $result->isNeutral(); // false
Export the result
Turn the result into an array, or encode it as JSON.
$result->toArray(); // ['label' => 'positive', 'compound' => 0.6588, // 'positive' => 0.594, 'negative' => 0.0, // 'neutral' => 0.406] json_encode($result); // {"label":"positive","compound":0.6588, // "positive":0.594,"negative":0, // "neutral":0.406}
Analyze Indonesian
Pass the language as an enum case, or as the code 'id'. An unknown code throws a ValueError.
use Risan\Sentiment\Language; use Risan\Sentiment\Sentiment; $result = Sentiment::analyze( 'Filmnya bagus banget!', Language::Indonesian, ); $result->label; // Label::Positive $result->compound; // 0.623 Sentiment::analyze('Filmnya bagus banget!', 'id');
Customize the analyzer
Add or remove words and change the threshold. Every with*() call returns a new analyzer.
use Risan\Sentiment\Analyzer; use Risan\Sentiment\Language; $analyzer = (new Analyzer(Language::Indonesian)) ->withWords(['cuan' => 2.5]) ->withoutWords(['kasar']) ->withThreshold(0.1); $result = $analyzer->analyze('Investasinya cuan!'); $result->label; // Label::Positive $result->compound; // 0.5848
The default threshold is 0.05. A text is positive when compound is at least the threshold. It is negative when compound is at most minus the threshold. Otherwise it is neutral.
More in the documentation: understanding the scores, languages, customizing the lexicon, thresholds and long text.
How it works
The engine is a PHP version of VADER (Hutto & Gilbert, 2014).
- A lexicon gives each known word a score from -4 to +4.
- Rules adjust the scores in context: negation, intensifiers and dampeners ("very", "kind of"), the contrast word "but", ALL CAPS,
!and?, emoticons and emoji. - The package adds up the scores and squashes the sum into
compound, between -1 and +1.
English uses the original VADER lexicon and rules. Its scores match vaderSentiment 3.3.2 on more than 2,700 reference scores, with two deliberate fixes. Indonesian runs the same engine with a lexicon and rules written for this project. It handles negation, intensifiers before and after the word (bagus banget), contrast words (tapi), informal spellings and emoji.
It is a lexicon method. It does not understand sarcasm, regional languages or words from your own field. How it works, in the docs.
Accuracy and performance
On the test split of IndoNLU SmSA, Indonesian scoring reaches 81.2% accuracy and 0.744 macro F1. The split has 500 reviews and comments in three classes, and the test uses the default threshold. Always guessing the most common class gives 41.6% accuracy. A fine-tuned model reaches more. Reproduce the numbers with php tools/evaluate-id.php --split=test. See Languages.
With OPcache on, the package scores about 80,000 short English texts per second and about 14,000 reviews per second. These are single runs on PHP 8.5, on a laptop with Docker on WSL2. They vary by 10% or more. The whole benchmark process peaks at about 2 MB. Run php benchmarks/run.php on your own hardware. See Performance.
Development
composer install composer test # Pest composer lint # Pint (PER style), check only composer analyse # PHPStan, level max composer bench # throughput and memory composer build:lexicons # regenerate resources/ from tools/data/
resources/ is generated from tools/data/. Run composer build:lexicons after you edit the sources. A test fails when they drift apart.
The examples in this README, on the website and in the docs are checked by tests/Feature/DocsExamplesTest.php.
Credits and license
MIT, see LICENSE.md. Third-party notices, including VADER's license and citation, are in NOTICE.md.