dorsal/arxiv
View SchemaReconsidering the significance of genomic word frequency
| Authors | Miklós Csűrös, Laurent Noé, Gregory Kucherov |
|---|---|
| Categories | |
| ArXiv ID | q-bio/0609022 |
| URL | https://arxiv.org/abs/q-bio/0609022 |
Abstract
We propose that the distribution of DNA words in genomic sequences can be primarily characterized by a double Pareto-lognormal distribution, which explains lognormal and power-law features found across all known genomes. Such a distribution may be the result of completely random sequence evolution by duplication processes. The parametrization of genomic word frequencies allows for an assessment of significance for frequent or rare sequence motifs.
{
"annotation_id": "c19a9f95-bb3b-4b0d-bca2-344719ccff39",
"date_created": "2026-03-02T18:01:35.333000Z",
"date_modified": "2026-03-02T18:01:35.333000Z",
"file_hash": "5cb4c1fa7bc8c658d9345f598d2d12cb87867f570ffa03f2825266ea14c343b3",
"private": false,
"record": {
"abstract": "We propose that the distribution of DNA words in genomic sequences can be\nprimarily characterized by a double Pareto-lognormal distribution, which\nexplains lognormal and power-law features found across all known genomes. Such\na distribution may be the result of completely random sequence evolution by\nduplication processes. The parametrization of genomic word frequencies allows\nfor an assessment of significance for frequent or rare sequence motifs.",
"arxiv_id": "q-bio/0609022",
"authors": [
"Mikl\u00f3s Cs\u0171r\u00f6s",
"Laurent No\u00e9",
"Gregory Kucherov"
],
"categories": [
"q-bio.GN"
],
"title": "Reconsidering the significance of genomic word frequency",
"url": "https://arxiv.org/abs/q-bio/0609022"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "3be1b665-2a08-47b8-a43b-ff6d7f02d45e",
"id": "arXiv Dataset IDs",
"type": "Model",
"variant": "snapshot-2026-03-01",
"version": "0.1.0"
},
"user_id": 1000002
}