fastText
| fastText | |
|---|---|
| Developer | Facebook's AI Research (FAIR) lab[1] |
| Release | November 9, 2015 |
| Stable release | 0.9.2[2]
/ April 28, 2020 |
| Written in | C++, Python |
| Platform | Linux, macOS, Windows |
| Type | Machine learning library |
| License | MIT License |
| Website | fasttext |
| Repository | github |
fastText is an open-source library developed by Facebook AI Research (FAIR) for learning word representations and performing text classification. It was publicly released in 2016.[3]
For word representation, fastText incorporates subword information using character n-grams, which enables it to construct representations for words not encountered during training.[4] The library also provides supervised methods designed for efficient text classification.[5]
The GitHub repository was archived on March 19, 2024.
Word representations
[edit]fastText builds on the skip-gram model used in word2vec, but also takes the internal structure of words into account. Instead of learning a representation for each word only as a whole, it represents words using character n-grams. Words that share character sequences can therefore share some of the same learned information.[6]
This use of subword information is particularly useful for rare words and languages with complex word formation, and also allows fastText to construct representations for words that were not seen during training.[6] Facebook AI Research later released pretrained fastText vectors for many languages, including a collection for 157 languages trained on Common Crawl and Wikipedia.[7]
Text classification
[edit]fastText can also be used for supervised text classification. It combines information from the words and word n-grams in a text to predict a class label. The method was designed to be computationally efficient and, in its original evaluation, achieved accuracy comparable to several contemporary deep-learning models while training and making predictions much faster.[8]
See also
[edit]References
[edit]- ↑ Mannes, John. "Facebook's fastText library is now optimized for mobile". TechCrunch. Retrieved 12 January 2018.
- ↑ Onur Çelebi (2020-04-28). "facebookresearch/fastText/releases/tag/v0.9.2". Facebook. Retrieved 2020-11-21.
- ↑ "Releasing fastText". fastText. 18 August 2016. Retrieved 16 August 2026.
- ↑ Bojanowski, Piotr; Grave, Edouard; Joulin, Armand; Mikolov, Tomas (2017). "Enriching Word Vectors with Subword Information". Transactions of the Association for Computational Linguistics. 5: 135–146. doi:10.1162/tacl_a_00051.
- ↑ Joulin, Armand; Grave, Edouard; Bojanowski, Piotr; Mikolov, Tomas (2017). "Bag of Tricks for Efficient Text Classification". Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. Association for Computational Linguistics. pp. 427–431.
- 1 2 Bojanowski, Piotr; Grave, Edouard; Joulin, Armand; Mikolov, Tomas (2017). "Enriching Word Vectors with Subword Information". Transactions of the Association for Computational Linguistics. 5: 135–146. doi:10.1162/tacl_a_00051.
- ↑ Grave, Edouard; Bojanowski, Piotr; Gupta, Prakhar; Joulin, Armand; Mikolov, Tomas (2018). "Learning Word Vectors for 157 Languages". Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). European Language Resources Association.
- ↑ Joulin, Armand; Grave, Edouard; Bojanowski, Piotr; Mikolov, Tomas (2017). "Bag of Tricks for Efficient Text Classification". Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. Association for Computational Linguistics. pp. 427–431.
External links
[edit]- fastText
- "FastText - Facebook Research Downloads". Archived from the original on 2020-11-03.
