
Sold by: Cotonoha
Open data
|
Deployed on AWS
Japanese Tokenizer Dictionaries for use with MeCab.
Overview
Japanese Tokenizer Dictionaries for use with MeCab.
Features and programs
Open Data Sponsorship Program
This dataset is part of the Open Data Sponsorship Program, an AWS program that covers the cost of storage for publicly available high-value cloud-optimized datasets.
Pricing
This is a publicly available data set. No subscription is required.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Legal
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
Delivery details
AWS Data Exchange (ADX)
AWS Data Exchange is a service that helps AWS easily share and manage data entitlements from other organizations at scale.
Open data resources
Available with or without an AWS account.
- How to use
- To access these resources, reference the Amazon Resource Name (ARN) using the AWS Command Line Interface (CLI). Learn more
- Description
- Dictionary Files
- Resource type
- S3 bucket
- Amazon Resource Name (ARN)
- arn:aws:s3:::cotonoha-dic
- AWS region
- ap-northeast-1
- AWS CLI access (No AWS account required)
- aws s3 ls --no-sign-request s3://cotonoha-dic/
Resources
Vendor resources
Support
Contact
Managed By
Cotonoha
How to cite
Japanese Tokenizer Dictionaries was accessed on DATE from https://registry.opendata.aws/cotonoha-dic .
License
Versions of Unidic offered here are available under the GPL/LGPL/BSD license.
IPADic is offered under a unique BSD-like license. See below.
<https://github.com/polm/ipadic-py/blob/master/ipadic/dicdir/COPYING>Similar products
YomiToku is a proprietary document analysis engine specialized for Japanese. It integrates AI OCR plus layout and table parsing models, accurately structuring vertical text, multi column documents, and complex business forms. It supports a wide range of use cases, including generating data for RAG / search, creating searchable PDFs, and extracting information from table data.
We provide yomitoku-client as a client SDK to help you use this product more conveniently.For more details, please refer to the link below:
https://github.com/MLism-Inc/yomitoku-client
For long-term or large-scale use, this product is also available through private offers.Please contact our support team for pricing information.

mixture of experts (MoE) architecture that supports multiple languages, primarily English and Japanese

Multilingual text question-answering retrieval, transforming textual information into dense vector representations.
