Documentation

Proper documentation is available at https://malaysian-dataset.readthedocs.io

How we gather dataset?

Crawling

Contributors heavily crawled Malaysian websites, you can check out the full list of crawled websites at https://github.com/users/huseinzol05/projects/1

Social media

We catch most of live data from Twitter, Facebook and Instagram using crawlers, So we just search using Elasticsearch query.

Translation

We use Google Translate.
We use LLM, including ChatGPT3.5, ChatGPT4, Mixtral, LLama3 70B.
We use Malaya translation, https://huggingface.co/mesolitica/translation-t5-small-standard-bahasa-cased-v2

Semisupervised

Teacher-student

Supervised small samples and then trained a base model.
Trained base model predict larger samples, retrain next student models on high confident labelled data.
Repeat.

LLM

Generate using ChatGPT3.5, ChatGPT4, Mixtral, LLama3 70B.

Notes

Any missing mp.py, get it at https://gist.github.com/huseinzol05/98974ae8c6c7a65d4bc0af9f5003786a
Any missing python scripts, please contact me ASAP or create an issue.
Please at least email us first before distributing these data. Remember all these hard workings we want to give it for free.
What do you see just the data, but nobody can see how much we spent our cost to make it public.

Suggestion

Feel free to contact me to request new dataset.
Feel free to open an issue if the link to dataset is forbidden, sometime I forgot to make it open to public.

Non-commercial Usage

A lot of data here semisupervised / translated / tagged / decoded using third party software, example, Google Translate, Google Speech, so to avoid any future complication, it is better not use this data for commercial purposes but allow for certain research purposes.

Acknowledgement

Thanks to Im Big, LigBlou, Mesolitica and KeyReply for sponsoring AWS Google and private cloud to deploy distributed crawlers.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

README.rst

README.rst

Documentation

How we gather dataset?

Crawling

Social media

Translation

Semisupervised

Teacher-student

LLM

Notes

Suggestion

Non-commercial Usage

Acknowledgement

Files

README.rst

Latest commit

History

README.rst

File metadata and controls

Documentation

How we gather dataset?

Crawling

Social media

Translation

Semisupervised

Teacher-student

LLM

Notes

Suggestion

Non-commercial Usage

Acknowledgement