DataHub: A Task-First Data Infrastructure for Latin American AI
A recent publication on arXiv introduces DataHub, a data infrastructure aimed at filling the gap in the dataset layer crucial for AI progress in Latin America. The paper highlights two key deficiencies in the region's AI framework: the dataset layer and the benchmark layer, prioritizing the former. The dataset layer encounters two main challenges: discovery and availability. Although Latin American AI datasets are present, they are dispersed across various platforms without a unified index, complicating access. Even with optimal indexing, the overall data volume is insufficient for advanced AI development. DataHub is suggested as a task-oriented data infrastructure structured by the ontology /<task?>/<domain?>/<language?>, incorporating features for dataset discovery, metadata, contribution, licensing, and reuse. This work falls under Computer Science > Artificial Intelligence and is recorded on arXiv with the identifier 2608.02949. The initiative seeks to establish a centralized and standardized platform to enhance dataset accessibility and supply, potentially fostering stronger AI development in Latin America.
Key facts
- Latin America lacks two foundational layers of AI infrastructure: dataset and benchmark layers.
- The paper focuses on the dataset layer.
- The dataset layer faces problems of discovery and supply.
- Latin American AI datasets are scattered across platforms with no shared index.
- Even with perfect indexing, total data volume is insufficient for frontier AI development.
- DataHub is a task-first data infrastructure.
- DataHub is organized through the ontology /<task?>/<domain?>/<language?>/.
- DataHub includes mechanisms for dataset discovery, metadata, contribution, licensing, and reuse.
- The paper is categorized under Computer Science > Artificial Intelligence.
- The paper was submitted to arXiv with identifier 2608.02949.
Entities
Institutions
- arXiv
Locations
- Latin America