Conversation
…n bwtdisk_prepare.
|
I'd suggest picking the lexicographically smallest possible nucleotide for that ambiguity code rather than a random one, to make the result deterministic. |
|
I have updated the function so that the lexicographically smallest possible nucleotide for each ambiguity code. Thanks for the suggestion @sjackman |
| } | ||
| // get the lexicographically smallest base for the code | ||
| char base = IUPAC::getPossibleSymbols(line[pos])[0]; | ||
| line[pos] = base; |
There was a problem hiding this comment.
You can omit the base intermediate variable.
line[pos] = IUPAC::getPossibleSymbols(line[pos])[0];|
What's your use case, Cole? Is it that you have reads with |
|
My use case is using assembled genomes. Ideally, I would like to be able to keep Do you know how hard it would be to accept |
|
Jared (@jts) is in a better position to answer that question than myself. |
These changes adds functionality to accept ambiguous IUPAC codes in
bwtdisk_prepare.It handles the ambiguity by choosing a random base that is within the set of bases for that ambiguity code. For example,
Ncan beA,C,G, orT;Scan beCorG;Hcan beA,C, orT; etc.This required
libdbgfmto be linked tobwtdisk_preparebecause it uses the IUPAC methods found inalphabet.cpp.